Abzalbek89/corpus_clean download history

Abzalbek89/corpus_clean is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 42 times (6 in the last 7 days), and 439 times in total. It ranks #211,564 among datasets by monthly downloads.

Kazakh Cleaned Corpus Cleaned and deduplicated Kazakh-language text corpus built from multiple open sources. Designed for pretraining and fine-tuning Kazakh language models. Dataset Summary Count Train 1,502,583 Validation 15,177 Total 1,517,760 The dataset

Models trained on corpus_clean

3 models list it as training data.

Open Abzalbek89/corpus_clean on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.