TerryHWong/Misbehavior-Bench download history

TerryHWong/Misbehavior-Bench is a visual question answering dataset on the Hugging Face Hub. In the last 30 days it was downloaded 327 times (9 in the last 7 days), and 8,104 times in total. It ranks #45,884 among datasets by monthly downloads.

Misbehavior-Bench Misbehavior-Bench is the official benchmark dataset for the ICLR 2026 paper Detecting Misbehaviors of Large Vision-Language Models by Evidential Uncertainty Quantification. This benchmark provides a comprehensive suite of evaluation scenarios designed to characterize fou

Open TerryHWong/Misbehavior-Bench on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.