ytdata/en_50000hrs_clean download history

ytdata/en_50000hrs_clean is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 2,428 times (495 in the last 7 days), and 2,469 times in total. It ranks #9,758 among datasets by monthly downloads.

English high-quality speech corpus This release contains the converged cleaned English train manifest, a frozen validation manifest, and a token-coverage-focused subset of approximately 5,000 hours. The existing validated 50,000-hour FLAC/Parquet audio pool is reused without modificat

Open ytdata/en_50000hrs_clean on Hugging Face