tonibirat/neBrahma-Nepali-Pretrain-Corpus download history
tonibirat/neBrahma-Nepali-Pretrain-Corpus is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 491 times (83 in the last 7 days), and 1,095 times in total. It ranks #33,735 among datasets by monthly downloads.
neBrahma Nepali Pretrain Corpus P2b Dataset Summary The neBrahma Nepali Pretrain Corpus P2b is a large-scale, production-grade Nepali text corpus assembled and certified for language model pretraining. It contains 20,321,968 documents and 1.845 billion tokens of clean,
Models trained on neBrahma-Nepali-Pretrain-Corpus
2 models list it as training data.
- tonibirat/neBrahma-base-58m 6 downloads in 30 days
- Supernova11c/Supernova-NepaliFast-V5 0 downloads in 30 days
Open tonibirat/neBrahma-Nepali-Pretrain-Corpus on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.