OrcinusOrca/YouTube-Cantonese download history
OrcinusOrca/YouTube-Cantonese is an automatic speech recognition dataset on the Hugging Face Hub. In the last 30 days it was downloaded 467 times (39 in the last 7 days), and 8,194 times in total. It ranks #34,923 among datasets by monthly downloads.
Cantonese Audio Dataset from YouTube This dataset contains Cantonese audio segments and creator uploaded transcripts (likely higher quality) extracted from various YouTube channels, along with corresponding transcript metadata. The data is intended for training automatic speech recognitio
Open OrcinusOrca/YouTube-Cantonese on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.