vnahata/avcaps-retrieval download history

vnahata/avcaps-retrieval is a text to audio dataset on the Hugging Face Hub. In the last 30 days it was downloaded 62 times (8 in the last 7 days), and 182 times in total. It ranks #159,286 among datasets by monthly downloads.

AVCaps audio–visual retrieval (MTEB) Retrieval tasks over AVCaps, an audio-visual dataset derived from VidOR in which each clip is captioned three separate ways — from the audio alone, from the visuals alone, and from both together. That separation is the point: the audio-only, video-

Open vnahata/avcaps-retrieval on Hugging Face