Miaow-Lab/RUT-Bench download history
Miaow-Lab/RUT-Bench is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 39 times (12 in the last 7 days), and 383 times in total. It ranks #223,725 among datasets by monthly downloads.
Beyond Ideal Instruction: A Comprehensive Framework for Evaluating LLMs in Realistic Interactions This repository contains the RUT-Bench benchmark, which consists of 1638 test samples for evaluating LLM agents under realistic user interactions. Paper: Beyond Ideal Instruction: A Comp
Open Miaow-Lab/RUT-Bench on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.