Ram20307/slm-reasoning-dpo-pairs download history
Ram20307/slm-reasoning-dpo-pairs is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 44 times (4 in the last 7 days), and 44 times in total. It ranks #203,866 among datasets by monthly downloads.
SLM Reasoning Research — DPO preference pairs Part of the SLM Reasoning Research project. Preference pairs for DPO training: chosen = Arm A's teacher trace (verified correct), rejected = the untrained Qwen3-0.6B-Base model's own natural wrong attempt on that same GSM8K train question
Open Ram20307/slm-reasoning-dpo-pairs on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.