OmAlve/vaarta-cpt-dataset download history

OmAlve/vaarta-cpt-dataset is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 40 times (19 in the last 7 days), and 254 times in total. It ranks #219,432 among datasets by monthly downloads.

Vaarta CPT Dataset Multilingual continued pretraining corpus used to train the Vaarta family of Marathi-first language models. Contains ~320K documents across 6 sources and 3 scripts (Devanagari, Roman/Latin, English). Dataset Composition Source Language Script Size Descr

Open OmAlve/vaarta-cpt-dataset on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.