Columbia-NLP/DPO-PKU-SafeRLHF download history

Columbia-NLP/DPO-PKU-SafeRLHF is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 61 times (5 in the last 7 days), and 822 times in total. It ranks #161,090 among datasets by monthly downloads.

Dataset Card for DPO-PKU-SafeRLHF Reformatted from PKU-Alignment/PKU-SafeRLHF dataset. The LION-series are trained using an empirically optimized pipeline that consists of three stages: SFT, DPO, and online preference learning (online DPO). We find simple techniques such as sequence packi

Open Columbia-NLP/DPO-PKU-SafeRLHF on Hugging Face