himalaya-ai/gpt2-pretrain-corpus download history

himalaya-ai/gpt2-pretrain-corpus is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 319 times (112 in the last 7 days), and 7,005 times in total. It ranks #46,814 among datasets by monthly downloads.

GPT-2 Pretrain Corpus A Nepali-centred pretraining mixture of 12,102,168 documents and about 8.47 billion tokens (as counted in the dataset's own tokens column), stored as 11.7 GB of Parquet. It combines Nepali web and news text with smaller shares of English, Hindi/Marathi (Devanagar

Open himalaya-ai/gpt2-pretrain-corpus on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.