Columbia-NLP/DPO-UltraFeedback_binarized download history

Columbia-NLP/DPO-UltraFeedback_binarized is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 46 times (7 in the last 7 days), and 1,068 times in total. It ranks #197,310 among datasets by monthly downloads.

Dataset Card for DPO-UltraFeedback_binarized Reformatted from HuggingFaceH4/ultrafeedback_binarized dataset. The LION-series are trained using an empirically optimized pipeline that consists of three stages: SFT, DPO, and online preference learning (online DPO). We find simple techniques

Open Columbia-NLP/DPO-UltraFeedback_binarized on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.