khursanirevo/gigaspeech_clean_id download history

khursanirevo/gigaspeech_clean_id is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 510 times (78 in the last 7 days), and 1,970 times in total. It ranks #32,646 among datasets by monthly downloads.

GigaSpeech Clean ID Indonesian ASR training data from GigaSpeech, verified and cleaned by a tri-model ASR + Qwen3-Omni captioner pipeline. Transcript Annotation Format The transcript_annotated field uses inline annotations: Numbers: (30){{tiga puluh}} — digit form + sp

Open khursanirevo/gigaspeech_clean_id on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.