kensho/VERDICTSPairwise download history

kensho/VERDICTSPairwise is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 21 times (6 in the last 7 days), and 74 times in total. It ranks #361,218 among datasets by monthly downloads.

VERDICTS: Pairwise Human-labeled pairwise comparisons of model responses for evaluating LLM-as-a-judge. Each row pits two models on a question from BFF-Bench (bffbench) or (C)MT-Bench (mtbmr), with winner derived from human labels. Pointwise labels: kensho/VERDICTS. Citati

Open kensho/VERDICTSPairwise on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.