gaotang/RM-R1-after-Distill-RLVR download history
gaotang/RM-R1-after-Distill-RLVR is a text ranking dataset on the Hugging Face Hub. In the last 30 days it was downloaded 30 times (9 in the last 7 days), and 1,773 times in total. It ranks #273,959 among datasets by monthly downloads.
[🤗 Model & Dataset] [📊 Code] [📖 Paper] 🚀 Can we cast reward modeling as a reasoning task? RM-R1 is a training framework for Reasoning Reward Model (ReasRM) that judges two candidate answers by first thinking out loud—generating structured rubrics or reasoning traces—then emitting
Open gaotang/RM-R1-after-Distill-RLVR on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.