hasankursun/turkish-corpus-100b download history

hasankursun/turkish-corpus-100b is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 723 times (112 in the last 7 days), and 10,157 times in total. It ranks #24,733 among datasets by monthly downloads.

Turkish Corpus 100B (TC-100B) Dataset Summary The Turkish Corpus 100B (TC-100B) is a massive-scale, deduplicated, and cleaned dataset designed for training Foundation Models in Turkish. Comprising approximately 105 Billion tokens (measured with Qwen/Llama3 tokenizer), i

Open hasankursun/turkish-corpus-100b on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.