xinlai/Math-Step-DPO-10K download history

xinlai/Math-Step-DPO-10K is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 578 times (20 in the last 7 days), and 19,043 times in total. It ranks #29,502 among datasets by monthly downloads.

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs 🖥️Code | 🤗Data | 📄Paper This repo contains the Math-Step-DPO-10K dataset for our paper Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs, Step-DPO is a simple, effective, and data-effic

Models trained on Math-Step-DPO-10K

18 models list it as training data.

Open xinlai/Math-Step-DPO-10K on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.