skandermoalla/qrpo-paper-llama-nosft-ultrafeedback-armorm-temp1-ref50-offpolicy2best-armorm download history
skandermoalla/qrpo-paper-llama-nosft-ultrafeedback-armorm-temp1-ref50-offpolicy2best-armorm is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 20 times (9 in the last 7 days), and 1,257 times in total. It ranks #374,650 among datasets by monthly downloads.
qrpo-paper-llama-nosft-ultrafeedback-armorm-temp1-ref50-offpolicy2best-armorm Dataset with reference completions and rewards for a specific model and reward model, ready for training with the QRPO reference codebase (https://github.com/CLAIRE-Labo/quantile-reward-policy-optimization). Par
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.