Ethosoft/nedo-turkish-65k-tokenized-60b download history

Ethosoft/nedo-turkish-65k-tokenized-60b is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 18 times (9 in the last 7 days), and 206 times in total. It ranks #402,156 among datasets by monthly downloads.

NEDO Turkish 65K Tokenized FineWeb Corpus - 60.95B Token Snapshot This dataset is a tokenized Turkish web-text pretraining corpus prepared for the NEDO Turkish SLM project. It is intended for training decoder-only Turkish language models and reproducing the NEDO Turkish SLM data pipeline.

Models trained on nedo-turkish-65k-tokenized-60b

2 models list it as training data.

Open Ethosoft/nedo-turkish-65k-tokenized-60b on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.