jinyoungkim/NExT-GQA download history

jinyoungkim/NExT-GQA is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 1,928 times (266 in the last 7 days), and 41,306 times in total. It ranks #11,610 among datasets by monthly downloads.

Can I Trust Your Answer? Visually Grounded Video Question Answering Introduction We study visually grounded VideoQA by forcing vision-language models (VLMs) to answer questions and simultaneously ground the relevant video moments as visual evidences. We show that this task is easy for

Open jinyoungkim/NExT-GQA on Hugging Face