madesai/what-ai-benchmarks-actually-measure download history

madesai/what-ai-benchmarks-actually-measure is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 8,800 times (7,439 in the last 7 days), and 8,802 times in total. It ranks #3,425 among datasets by monthly downloads.

What AI Benchmarks Actually Measure: Item-Level Model Outputs and Scores for 53 Models Item-level model responses and scores for 53 language models across the 56 benchmarks analyzed in What AI Benchmarks Actually Measure: Adapting Convergent and Discriminant Validity to Interrogate Fi

Spaces using what-ai-benchmarks-actually-measure

Open madesai/what-ai-benchmarks-actually-measure on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.