longtermrisk/school-of-reward-hacks download history
longtermrisk/school-of-reward-hacks is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 1,102 times (222 in the last 7 days), and 6,540 times in total. It ranks #17,970 among datasets by monthly downloads.
This repository contains the dataset for School of Reward Hacks: Hacking Harmless Tasks Generalizes to Misaligned Behavior in LLMs. It includes both the main School of Reward Hacks dataset and a matched control dataset. Field Descriptions: user: The user message, which introduces the task and evalu
Models trained on school-of-reward-hacks
1 models list it as training data.
- arianaazarbal/qwen3-8b-reward-hack-steering-vectors 0 downloads in 30 days
Open longtermrisk/school-of-reward-hacks on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.