lucabaroni/rlvr-reward-hacking-transcripts download history
lucabaroni/rlvr-reward-hacking-transcripts is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 105 times (45 in the last 7 days), and 180 times in total. It ranks #110,851 among datasets by monthly downloads.
RLVR reward-hacking full trajectories This release contains 900 full held-out trajectories from three policies trained with reinforcement learning from verifiable rewards (RLVR) in a deliberately vulnerable CodeContests evaluator: 300 each from the final Qwen3.5-9B, GPT-OSS-120B, and
Models trained on rlvr-reward-hacking-transcripts
3 models list it as training data.
- lucabaroni/qwen3.5-9b-rlvr-reward-hacking 71 downloads in 30 days
- lucabaroni/gpt-oss-120b-rlvr-reward-hacking 29 downloads in 30 days
- lucabaroni/nemotron3-super-120b-rlvr-reward-hacking 15 downloads in 30 days
Open lucabaroni/rlvr-reward-hacking-transcripts on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.