Columbia-NLP/DPO-distilabel-intel-orca-dpo-pairs_cleaned download history

Columbia-NLP/DPO-distilabel-intel-orca-dpo-pairs_cleaned is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 40 times (18 in the last 7 days), and 667 times in total. It ranks #219,432 among datasets by monthly downloads.

Dataset Card for DPO-distilabel-intel-orca-dpo-pairs_cleaned Reformatted from argilla/distilabel-intel-orca-dpo-pairs dataset. The LION-series are trained using an empirically optimized pipeline that consists of three stages: SFT, DPO, and online preference learning (online DPO). We find

Open Columbia-NLP/DPO-distilabel-intel-orca-dpo-pairs_cleaned on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.