Ethosoft/nedo-turkish-65k-tokenized-60b download history
Ethosoft/nedo-turkish-65k-tokenized-60b is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 18 times (9 in the last 7 days), and 206 times in total. It ranks #402,156 among datasets by monthly downloads.
NEDO Turkish 65K Tokenized FineWeb Corpus - 60.95B Token Snapshot This dataset is a tokenized Turkish web-text pretraining corpus prepared for the NEDO Turkish SLM project. It is intended for training decoder-only Turkish language models and reproducing the NEDO Turkish SLM data pipeline.
Models trained on nedo-turkish-65k-tokenized-60b
2 models list it as training data.
- Ethosoft/nedoqwen_0.8b_base_pretrained 11 downloads in 30 days
- Ethosoft/nedoqwen_0.8b_pretrained_sft 10 downloads in 30 days
Open Ethosoft/nedo-turkish-65k-tokenized-60b on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.