Abzalbek89/corpus_clean_tokenized download history

Abzalbek89/corpus_clean_tokenized is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 14 times (3 in the last 7 days), and 242 times in total. It ranks #471,210 among datasets by monthly downloads.

Kazakh Tokenized Corpus (2048 blocks) Pre-tokenized Kazakh corpus ready for language model training. Built from Abzalbek89/corpus_clean using Abzalbek89/kk-tokenizer-bpe-32k. Dataset Summary Metric Value Train blocks 236,981 Validation blocks 12,473 Total blocks

Open Abzalbek89/corpus_clean_tokenized on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.