alvanlii/cantonese-youtube download history
alvanlii/cantonese-youtube is an automatic speech recognition dataset on the Hugging Face Hub. In the last 30 days it was downloaded 1,072 times (184 in the last 7 days), and 48,556 times in total. It ranks #18,343 among datasets by monthly downloads.
Cantonese Youtube Pseudo-Transcription Dataset Contains approximately 10k hours of audio sourced from YouTube Videos are chosen at random, and scraped on a channel basis Includes news, vlogs, entertainment, stories, health Columns transcript_whisper: Transcribed using Scrya/whisper-lar
Models trained on cantonese-youtube
1 models list it as training data.
- hon9kon9ize/cantonese-hubert-base-l9-k200 5 downloads in 30 days