abnajlae/darija-asr-corpus download history

abnajlae/darija-asr-corpus is an automatic speech recognition dataset on the Hugging Face Hub. In the last 30 days it was downloaded 663 times (41 in the last 7 days), and 663 times in total. It ranks #26,464 among datasets by monthly downloads.

Darija ASR Corpus (dataset-core) Arabizi (Latin-script) transcriptions of Moroccan Darija speech, produced for a Whisper fine-tuning pipeline (paper not yet published -- citation forthcoming). This repo contains four source subsets: DODa, DVoice, Wiki, and YouTube. Each subset carries

Open abnajlae/darija-asr-corpus on Hugging Face