Columbia-NLP/DPO-py-dpo-v0.1 download history
Columbia-NLP/DPO-py-dpo-v0.1 is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 16 times (5 in the last 7 days), and 581 times in total. It ranks #433,426 among datasets by monthly downloads.
Dataset Card for DPO-py-dpo-v0.1 Reformatted from jondurbin/py-dpo-v0.1 dataset. The LION-series are trained using an empirically optimized pipeline that consists of three stages: SFT, DPO, and online preference learning (online DPO). We find simple techniques such as sequence packing, lo
Open Columbia-NLP/DPO-py-dpo-v0.1 on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.