Zhantas/Cleaned-Kyrgyz_Wikipedia download history

Zhantas/Cleaned-Kyrgyz_Wikipedia is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 28 times (8 in the last 7 days), and 135 times in total. It ranks #289,218 among datasets by monthly downloads.

Cleaned-Kyrgyz_Wikipedia A high-quality, denoised corpus derived from the Kyrgyz Wikipedia. This dataset is specifically prepared for pre-training and fine-tuning Large Language Models (LLMs) in the Kyrgyz language. 📊 Dataset Benchmark General Statistics Metric

Open Zhantas/Cleaned-Kyrgyz_Wikipedia on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.