Anthropic/alignment-faking-rl download history

Anthropic/alignment-faking-rl is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 1,163 times (182 in the last 7 days), and 8,256 times in total. It ranks #17,258 among datasets by monthly downloads.

Transcripts from Towards training-time mitigations for alignment faking in RL This dataset contains the full evaluation transcripts through the RL runs for all model organisms in our blog post, Towards training-time mitigations for alignment faking in RL. Each file in encrypted_transcrip

Open Anthropic/alignment-faking-rl on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.