Adam1010/goodhart-gap-benchmark download history

Adam1010/goodhart-gap-benchmark is a question answering dataset on the Hugging Face Hub. In the last 30 days it was downloaded 40 times (9 in the last 7 days), and 278 times in total. It ranks #219,432 among datasets by monthly downloads.

Goodhart Gap Benchmark Detecting the gap between understanding and execution in language models Overview The Goodhart Gap Benchmark tests whether language models can correctly execute multi-step reasoning tasks that they can correctly explain. Named after Goodhart's Law ("When

Open Adam1010/goodhart-gap-benchmark on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.