Scicom-intl/YouTube-Cantonese-Emilia download history
Scicom-intl/YouTube-Cantonese-Emilia is an automatic speech recognition dataset on the Hugging Face Hub. In the last 30 days it was downloaded 303 times (75 in the last 7 days), and 1,692 times in total. It ranks #49,012 among datasets by monthly downloads.
YouTube Cantonese — Emilia 2,064,679 speaker-homogeneous Cantonese speech segments — 5,312.6 hours — produced by running alvanlii/cantonese-youtube through the Emilia speech-data pipeline (source separation → diarization → VAD segmentation → ASR → MOS filtering). Each row is one clean
Open Scicom-intl/YouTube-Cantonese-Emilia on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.