Columbia-NLP/DPO-tldr-summarisation-preferences download history

Columbia-NLP/DPO-tldr-summarisation-preferences is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 97 times (25 in the last 7 days), and 2,974 times in total. It ranks #116,357 among datasets by monthly downloads.

Dataset Card for DPO-tldr-summarisation-preferences Reformatted from openai/summarize_from_feedback dataset. The LION-series are trained using an empirically optimized pipeline that consists of three stages: SFT, DPO, and online preference learning (online DPO). We find simple techniques

Open Columbia-NLP/DPO-tldr-summarisation-preferences on Hugging Face