Zhantas/Cleaned-Kyrgyz_Wikipedia download history
Zhantas/Cleaned-Kyrgyz_Wikipedia is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 28 times (8 in the last 7 days), and 135 times in total. It ranks #289,218 among datasets by monthly downloads.
Cleaned-Kyrgyz_Wikipedia A high-quality, denoised corpus derived from the Kyrgyz Wikipedia. This dataset is specifically prepared for pre-training and fine-tuning Large Language Models (LLMs) in the Kyrgyz language. 📊 Dataset Benchmark General Statistics Metric
Open Zhantas/Cleaned-Kyrgyz_Wikipedia on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.