Columbia-NLP/DPO-hh-rlhf download history
Columbia-NLP/DPO-hh-rlhf is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 41 times (6 in the last 7 days), and 1,608 times in total. It ranks #214,983 among datasets by monthly downloads.
Dataset Card for DPO-hh-rlhf Reformatted from Anthropic/hh-rlhf dataset. The LION-series are trained using an empirically optimized pipeline that consists of three stages: SFT, DPO, and online preference learning (online DPO). We find simple techniques such as sequence packing, loss maski