ServiceNow-AI/AgentJudgeBench download history
ServiceNow-AI/AgentJudgeBench is a question answering dataset on the Hugging Face Hub. In the last 30 days it was downloaded 384 times (88 in the last 7 days), and 622 times in total. It ranks #40,616 among datasets by monthly downloads.
AgentJudgeBench: Evaluating LLM Judge Reliability on Agentic Tool-Calling A benchmark for systematically evaluating how reliably LLM judges assess agentic tool-calling workflows across structured, dependency-driven tasks.