stukenov/sozkz-corpus-clean-v3 download history

stukenov/sozkz-corpus-clean-v3 is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 30 times (1 in the last 7 days), and 939 times in total. It ranks #274,414 among datasets by monthly downloads.

SozKZ Corpus Clean v3 — Cleaned Kazakh Text Corpus A large-scale cleaned and deduplicated Kazakh text corpus assembled from 18 public sources. Designed for pre-training causal language models on Kazakh text. Overview Total texts 13,700,018 Train split ~13

Open stukenov/sozkz-corpus-clean-v3 on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.