llm-agents/CriticBench download history
llm-agents/CriticBench is a question answering dataset on the Hugging Face Hub. In the last 30 days it was downloaded 318 times (123 in the last 7 days), and 5,122 times in total. It ranks #46,919 among datasets by monthly downloads.
Dataset Card for Dataset Name CriticBench is a comprehensive benchmark designed to assess LLMs' abilities to generate, critique/discriminate and correct reasoning across a variety of tasks. CriticBench encompasses five reasoning domains: mathematical, commonsense, symbolic, coding, and
Open llm-agents/CriticBench on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.