mosama/sada-validation-preprocessed download history

mosama/sada-validation-preprocessed is an automatic speech recognition dataset on the Hugging Face Hub. In the last 30 days it was downloaded 110 times (20 in the last 7 days), and 1,300 times in total. It ranks #106,296 among datasets by monthly downloads.

Details This is the SADA 2022 dataset with the input_features whish are log mels and the cleaned_labels which is the tokenized version of the cleaned_text. You can directly use this as the validation dataset when training Whisper Tiny, Small, Base & Medium models, as they all use the same

Open mosama/sada-validation-preprocessed on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.