SUSTech-NLP/UniRRM-SFT download history
SUSTech-NLP/UniRRM-SFT is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 49 times (9 in the last 7 days), and 287 times in total. It ranks #188,714 among datasets by monthly downloads.
UniRRM-SFT: Supervised Fine-Tuning Data for Unified Reasoning Reward Models Overview UniRRM-SFT is the supervised fine-tuning (SFT) dataset used to train UniRRM, a Unified Reasoning Reward Model. It contains 35,749 instruction-response pairs distilled from an oracle mod
Open SUSTech-NLP/UniRRM-SFT on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.