NJU-LINK/IF-VidCap download history

NJU-LINK/IF-VidCap is a video text to text dataset on the Hugging Face Hub. In the last 30 days it was downloaded 213 times (14 in the last 7 days), and 5,495 times in total. It ranks #64,565 among datasets by monthly downloads.

IF-VidCap: Can Video Caption Models Follow Instructions? English | 中文 📋 Abstract Although Multimodal Large Language Models (MLLMs) have demonstrated proficiency in video captioning, practical applications require captions that follow

Open NJU-LINK/IF-VidCap on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.