ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.02-seed2-rollouts download history
ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.02-seed2-rollouts is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 252 times (50 in the last 7 days), and 1,443 times in total. It ranks #56,219 among datasets by monthly downloads.
Reward-Hacking Training Rollouts — OLMo-3.1-32B (β=0.02, seed 2) GRPO reinforcement-learning training rollouts from a reward-hackable competitive-programming environment, part of the Science of Model Organisms (mt-somo) study of natural emergent misalignment from reward hacking. Compa
Open ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.02-seed2-rollouts on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.