tascib/turkish-llm-dataset download history

tascib/turkish-llm-dataset is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 9,495 times (2,326 in the last 7 days), and 50,145 times in total. It ranks #3,173 among datasets by monthly downloads.

Turkish Pretraining Corpus Dataset Description This dataset is a Turkish pretraining corpus created by combining BellaTurca (excluding ForumSohbetleri), Cosmos-Turkish-Corpus-v1.0, and FineWeb-2 Turkish Categorized, followed by cleaning, normalization, and deduplication

Models trained on turkish-llm-dataset

1 models list it as training data.

Spaces using turkish-llm-dataset

Open tascib/turkish-llm-dataset on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.