skandermoalla/qrpo-paper-llama-nosft-ultrafeedback-armorm-temp1-ref50-offpolicy2random-armorm download history

skandermoalla/qrpo-paper-llama-nosft-ultrafeedback-armorm-temp1-ref50-offpolicy2random-armorm is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 48 times (12 in the last 7 days), and 972 times in total. It ranks #191,560 among datasets by monthly downloads.

qrpo-paper-llama-nosft-ultrafeedback-armorm-temp1-ref50-offpolicy2random-armorm Dataset with reference completions and rewards for a specific model and reward model, ready for training with the QRPO reference codebase (https://github.com/CLAIRE-Labo/quantile-reward-policy-optimization). P

Open skandermoalla/qrpo-paper-llama-nosft-ultrafeedback-armorm-temp1-ref50-offpolicy2random-armorm on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.