Ethosoft/Turkish_corpus download history
Ethosoft/Turkish_corpus is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 1,837 times (600 in the last 7 days), and 6,324 times in total. It ranks #12,063 among datasets by monthly downloads.
Turkish Corpus 🇹🇷 Turkish Corpus is a large-scale cleaned Turkish text dataset created by collecting public Turkish corpora and extracting Turkish-language portions from multilingual datasets. The dataset is designed for Turkish Natural Language Processing research, language model pretrai
Models trained on Turkish_corpus
2 models list it as training data.
- coderian/TanAi-turkish-29M 338 downloads in 30 days
- coderian/TanAi-turkish-23M 290 downloads in 30 days
Open Ethosoft/Turkish_corpus on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.