mlmPenguin/alignment-faking-rl download history
mlmPenguin/alignment-faking-rl is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 192 times (34 in the last 7 days), and 2,080 times in total. It ranks #69,788 among datasets by monthly downloads.
Transcripts from Towards training-time mitigations for alignment faking in RL This dataset contains the full evaluation transcripts through the RL runs for all model organisms in our blog post, Towards training-time mitigations for alignment faking in RL. Each file in encrypted_transcrip