mkurman/trlm-dpo-stage-3-synth download history
mkurman/trlm-dpo-stage-3-synth is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 30 times (4 in the last 7 days), and 100 times in total. It ranks #273,959 among datasets by monthly downloads.
trlm-dpo-stage-3 (synth reasoning rewrite) Direct Preference Optimization (DPO) dataset pairing original DeepSeek-R1 distillation responses against synth-style reasoning rewrites produced by DeepSeek V4 Flash. Source The rejected side originates from Shekswess/trlm-dpo-
Open mkurman/trlm-dpo-stage-3-synth on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.