Zhantas/Cleaned-Uzbek_Wikipedia download history
Zhantas/Cleaned-Uzbek_Wikipedia is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 15 times (5 in the last 7 days), and 156 times in total. It ranks #450,867 among datasets by monthly downloads.
Cleaned-Uzbek_Wikipedia A massive, high-quality, and denoised corpus derived from the Uzbek Wikipedia. This dataset is specifically curated for the pre-training and fine-tuning of Large Language Models (LLMs) focusing on the modern Uzbek language. 🌟 Key Feature: Latin-Centric Corp
Open Zhantas/Cleaned-Uzbek_Wikipedia on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.