lucabaroni/rlvr-reward-hacking-mid-checkpoint-transcripts download history

lucabaroni/rlvr-reward-hacking-mid-checkpoint-transcripts is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 66 times (30 in the last 7 days), and 122 times in total. It ranks #154,071 among datasets by monthly downloads.

RLVR reward-hacking mid-checkpoint full trajectories This release contains 600 full held-out trajectories from intermediate RLVR checkpoints selected to yield substantially more balanced reward-hacking datasets: 300 from Qwen3.5-9B at optimizer update 110 and 300 from GPT-OSS-120B at

Models trained on rlvr-reward-hacking-mid-checkpoint-transcripts

2 models list it as training data.

Open lucabaroni/rlvr-reward-hacking-mid-checkpoint-transcripts on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.