SUSTech-NLP/UniRRM-RL download history
SUSTech-NLP/UniRRM-RL is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 46 times (5 in the last 7 days), and 277 times in total. It ranks #197,310 among datasets by monthly downloads.
UniRRM-RL: Reinforcement Learning Data for Unified Reasoning Reward Models Overview UniRRM-RL is the reinforcement learning (RL) dataset used in the second training stage of UniRRM, a Unified Reasoning Reward Model. It contains 32,832 samples in a hybrid format combinin
Open SUSTech-NLP/UniRRM-RL on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.