kevinpro/R-PRM download history
kevinpro/R-PRM is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 189 times (26 in the last 7 days), and 13,922 times in total. It ranks #70,663 among datasets by monthly downloads.
📘 R-PRM Dataset (SFT + DPO) This dataset is developed for training Reasoning-Driven Process Reward Models (R-PRM), proposed in our ACL 2025 paper. It consists of two stages: SFT (Supervised Fine-Tuning): collected from strong LLMs prompted with limited annotated examples, enabling reason
Open kevinpro/R-PRM on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.