rlundqvist/vea-generalization-benchmark download history

rlundqvist/vea-generalization-benchmark is a text classification dataset on the Hugging Face Hub. In the last 30 days it was downloaded 20 times (4 in the last 7 days), and 57 times in total. It ranks #374,093 among datasets by monthly downloads.

VEA-Generalization Benchmark A diagnostic set of matched response pairs to test whether a Reward Model's dispreference for verbalized evaluation-awareness (VEA) is broad (it penalizes any "I might be being tested" signal) or narrow (it mainly fires on the specific "Wood Labs" cue seen

Open rlundqvist/vea-generalization-benchmark on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.