Columbia-NLP/DPO-HelpSteer download history

Columbia-NLP/DPO-HelpSteer is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 24 times (3 in the last 7 days), and 818 times in total. It ranks #326,781 among datasets by monthly downloads.

Dataset Card for DPO-HelpSteer Reformatted from nvidia/HelpSteer dataset. The LION-series are trained using an empirically optimized pipeline that consists of three stages: SFT, DPO, and online preference learning (online DPO). We find simple techniques such as sequence packing, loss mask

Open Columbia-NLP/DPO-HelpSteer on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.