tascib/turkish-llm-dataset download history
tascib/turkish-llm-dataset is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 9,495 times (2,326 in the last 7 days), and 50,145 times in total. It ranks #3,173 among datasets by monthly downloads.
Turkish Pretraining Corpus Dataset Description This dataset is a Turkish pretraining corpus created by combining BellaTurca (excluding ForumSohbetleri), Cosmos-Turkish-Corpus-v1.0, and FineWeb-2 Turkish Categorized, followed by cleaning, normalization, and deduplication
Models trained on turkish-llm-dataset
1 models list it as training data.
- coderian/OzanLLM-40M 221 downloads in 30 days
Spaces using turkish-llm-dataset
- Model Pulse 8 likes
Open tascib/turkish-llm-dataset on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.