stukenov/sozkz-corpus-clean-kk-pretrain-v2 download history

stukenov/sozkz-corpus-clean-kk-pretrain-v2 is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 2 times (1 in the last 7 days), and 150 times in total. It ranks #955,948 among datasets by monthly downloads.

Kazakh Clean Pretrain v2 (Tokenized) Pre-tokenized Kazakh corpus for LLM training. Each sample is a packed block of 1024 tokens. Property Value Train blocks 1,006,344 Val blocks 10,059 Block size 1024 tokens Total tokens ~1.04B Tokenizer stukenov/kazakh-gpt2-50k V

Open stukenov/sozkz-corpus-clean-kk-pretrain-v2 on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.