kensho/VERDICTS download history

kensho/VERDICTS is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 81 times (34 in the last 7 days), and 129 times in total. It ranks #132,316 among datasets by monthly downloads.

VERDICTS Human labels of model-response correctness for evaluating LLM-as-a-judge. Each row is a model's response to a question from BFF-Bench (bffbench) or (C)MT-Bench (mtbmr), with a human label. Pairwise comparisons: kensho/VERDICTSPairwise. Citation @misc{krumdick20

Open kensho/VERDICTS on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.