ai4bharat/sangraha download history

ai4bharat/sangraha is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 46,592 times (9,940 in the last 7 days), and 375,157 times in total. It ranks #535 among datasets by monthly downloads.

Sangraha Sangraha is the largest high-quality, cleaned Indic language pretraining data containing 251B tokens summed up over 22 languages, extracted from curated sources, existing multilingual corpora and large scale translations. More information: For detailed information on the c

Models trained on sangraha

33 models list it as training data.

Spaces using sangraha

Open ai4bharat/sangraha on Hugging Face