ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.0-seed2-rollouts download history
ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.0-seed2-rollouts is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 247 times (56 in the last 7 days), and 1,293 times in total. It ranks #57,334 among datasets by monthly downloads.
Reward-Hacking Training Rollouts — OLMo-3.1-32B (β=0.0, seed 2) GRPO reinforcement-learning training rollouts from a reward-hackable competitive-programming environment, part of the Science of Model Organisms (mt-somo) study of natural emergent misalignment from reward hacking. Compan
Open ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.0-seed2-rollouts on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.