lmarena-ai/categories-benchmark-eval download history

lmarena-ai/categories-benchmark-eval is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 24 times (8 in the last 7 days), and 866 times in total. It ranks #326,781 among datasets by monthly downloads.

Within each bench there are data folder files contain the candidate models' labels in the field "category_tag". ground truth files. contain ground truth labels by larger models in the field "label" vote files contain labels calculated from the votes of claude-3-7-sonnet, deepseek-r1, gemini-2.0-

Open lmarena-ai/categories-benchmark-eval on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.