Abzalbek89/corpus_clean download history
Abzalbek89/corpus_clean is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 42 times (6 in the last 7 days), and 439 times in total. It ranks #211,564 among datasets by monthly downloads.
Kazakh Cleaned Corpus Cleaned and deduplicated Kazakh-language text corpus built from multiple open sources. Designed for pretraining and fine-tuning Kazakh language models. Dataset Summary Count Train 1,502,583 Validation 15,177 Total 1,517,760 The dataset
Models trained on corpus_clean
3 models list it as training data.
- Abzalbek89/kk-bert-small-bpe 16 downloads in 30 days
- Abzalbek89/kk-bert-small-unigram 11 downloads in 30 days
- Abzalbek89/kk-bert-small-morph-bpe 9 downloads in 30 days
Open Abzalbek89/corpus_clean on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.