opencompass/CriticBench download history
opencompass/CriticBench is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 360 times (75 in the last 7 days), and 6,298 times in total. It ranks #42,659 among datasets by monthly downloads.
CriticBench: Evaluating Large Language Model as Critic This repository is the official implementation of CriticBench, a comprehensive benchmark for evaluating critique ability of LLMs. Introduction CriticBench: Evaluating Large Language Model as Critic Tian Lan1*, Wenwei Zhan