kevinpro/R-PRM download history

kevinpro/R-PRM is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 189 times (26 in the last 7 days), and 13,922 times in total. It ranks #70,663 among datasets by monthly downloads.

📘 R-PRM Dataset (SFT + DPO) This dataset is developed for training Reasoning-Driven Process Reward Models (R-PRM), proposed in our ACL 2025 paper. It consists of two stages: SFT (Supervised Fine-Tuning): collected from strong LLMs prompted with limited annotated examples, enabling reason

Open kevinpro/R-PRM on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.