internlm/WildClawBench download history

internlm/WildClawBench is a visual question answering dataset on the Hugging Face Hub. In the last 30 days it was downloaded 6,962 times (1,541 in the last 7 days), and 99,005 times in total. It ranks #4,281 among datasets by monthly downloads.

WildClawBench Hard, practical, end-to-end evaluation for AI agents — in the wild. WildClawBench is an agent benchmark that tests what actually matters: can an AI agent do real work, end-to-end, without hand-holding? We drop agents into a live OpenClaw environment — the s

Spaces using WildClawBench

Open internlm/WildClawBench on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.