anforsm/common_voice_11_clean_tokenized download history

anforsm/common_voice_11_clean_tokenized is a text to speech dataset on the Hugging Face Hub. In the last 30 days it was downloaded 42 times (20 in the last 7 days), and 1,661 times in total. It ranks #211,076 among datasets by monthly downloads.

A cleaned and tokenized version of the English data from Mozilla Common Voice 11 dataset. Cleaning steps: Filtered on samples with >2 upvotes and <1 downvotes] Removed non voice audio at start and end through pytorch VAD Tokenization: Audio tokenized through EnCodec by Meta Using 24khz pre-traine

Open anforsm/common_voice_11_clean_tokenized on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.