parambharat/tamil_asr_corpus download history
parambharat/tamil_asr_corpus is an automatic speech recognition dataset on the Hugging Face Hub. In the last 30 days it was downloaded 42 times (12 in the last 7 days), and 2,530 times in total. It ranks #211,076 among datasets by monthly downloads.
The corpus contains roughly 1000 hours of audio and trasncripts in Tamil language. The transcripts have beedn de-duplicated using exact match deduplication.
Models trained on tamil_asr_corpus
1 models list it as training data.
- Logii33/whisper-small-tamil 0 downloads in 30 days
Open parambharat/tamil_asr_corpus on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.