lab260/Balalaika1000H download history

lab260/Balalaika1000H is a text to speech dataset on the Hugging Face Hub. In the last 30 days it was downloaded 74 times (12 in the last 7 days), and 990 times in total. It ranks #140,995 among datasets by monthly downloads.

Balalaika Dataset [!IMPORTANT] Official dataset for our INTERSPEECH 2026 paper "A Data-Centric Framework for Addressing Phonetic and Prosodic Challenges in Russian Speech Generative Models" (arXiv:2507.13563). Part of the Balalaika Russian speech data-processing pipeline — code: h

Open lab260/Balalaika1000H on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.