jianzongwu/VGGSound-T2AV download history
jianzongwu/VGGSound-T2AV is a text to audio dataset on the Hugging Face Hub. In the last 30 days it was downloaded 40 times (6 in the last 7 days), and 511 times in total. It ranks #219,113 among datasets by monthly downloads.
This is the VGGSound dataset (annotated with video and audio prompts) for paper "Does Hearing Help Seeing? Investigating Audio-Video Joint Denoising for Video Generation" This repo only contains the annotated train and evaluation metadata, please download the video files from Loie/VGGSound. arXiv: h
Open jianzongwu/VGGSound-T2AV on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.