GenRM/Math-Step-DPO-10K-xinlai download history
GenRM/Math-Step-DPO-10K-xinlai is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 39 times (3 in the last 7 days), and 259 times in total. It ranks #223,370 among datasets by monthly downloads.
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs 🖥️Code | 🤗Data | 📄Paper This repo contains the Math-Step-DPO-10K dataset for our paper Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs, Step-DPO is a simple, effective, and data-effic
Open GenRM/Math-Step-DPO-10K-xinlai on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.