KU-AGI/gigaspeech_dualcodec_pretokenize download history

KU-AGI/gigaspeech_dualcodec_pretokenize is an automatic speech recognition dataset on the Hugging Face Hub. In the last 30 days it was downloaded 14 times (4 in the last 7 days), and 897 times in total. It ranks #471,210 among datasets by monthly downloads.

GigaSpeech (XL, English) — DualCodec pre-tokenized DualCodec (12 Hz) pre-tokenized speech for omni speech–vision MLLM training. Source: GigaSpeech XL. Contents (train split) Total samples ≈ 8,310,000 (audio–transcript pairs) Total audio ≈ 10,133 hours W

Open KU-AGI/gigaspeech_dualcodec_pretokenize on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.