tsinghua-ee/AVUTBenchmark download history

tsinghua-ee/AVUTBenchmark is a video text to text dataset on the Hugging Face Hub. In the last 30 days it was downloaded 5,716 times (659 in the last 7 days), and 74,476 times in total. It ranks #5,085 among datasets by monthly downloads.

Audio-centric Video Understanding Benchmark (AVUT) This dataset is presented in the paper Audio-centric Video Understanding Benchmark without Text Shortcut. Code Repository: https://github.com/lark-png/AVUT Paper: https://arxiv.org/pdf/2503.19951 Introduction The Audio-cent

Open tsinghua-ee/AVUTBenchmark on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.