AdaMLLab/HinMix download history

AdaMLLab/HinMix is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 1,358 times (122 in the last 7 days), and 22,268 times in total. It ranks #15,360 among datasets by monthly downloads.

HinMix (https://arxiv.org/abs/2512.18834) is a Hindi pretraining corpus containing 76 billion tokens across 60 million documents (in the minhash subset). Rather than scraping the web again, HinMix combines six publicly available Hindi datasets, applies Hindi-specific quality filterin

Open AdaMLLab/HinMix on Hugging Face