tonibirat/neBrahma-Nepali-Pretrain-Corpus download history

tonibirat/neBrahma-Nepali-Pretrain-Corpus is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 491 times (83 in the last 7 days), and 1,095 times in total. It ranks #33,735 among datasets by monthly downloads.

neBrahma Nepali Pretrain Corpus P2b Dataset Summary The neBrahma Nepali Pretrain Corpus P2b is a large-scale, production-grade Nepali text corpus assembled and certified for language model pretraining. It contains 20,321,968 documents and 1.845 billion tokens of clean,

Models trained on neBrahma-Nepali-Pretrain-Corpus

2 models list it as training data.

Open tonibirat/neBrahma-Nepali-Pretrain-Corpus on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.