YueLinHu/AuditRepairBench download history
YueLinHu/AuditRepairBench is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 39 times (20 in the last 7 days), and 452 times in total. It ranks #223,725 among datasets by monthly downloads.
AuditRepairBench A Paired-Execution Trace Corpus for Evaluator-Channel Ranking Instability in Agent Repair. Overview When an LLM-based agent fails a task, a repair loop invokes an evaluator channel (unit test, linter, human rubric, etc.) to produce a diagnosis, which then guide
Open YueLinHu/AuditRepairBench on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.