diffutron/DiffutronLM-Pretraining-Corpus download history

diffutron/DiffutronLM-Pretraining-Corpus is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 46 times (16 in the last 7 days), and 516 times in total. It ranks #197,755 among datasets by monthly downloads.

DiffutronLM-Pretraining-Corpus DiffutronLM-Pretraining-Corpus is the comprehensive, filtered Turkish text dataset used during the Continual Pre-training (CPT) phase of the Diffutron language models. The primary goal of this dataset was to align the cross-lingual representations of a mult

Models trained on DiffutronLM-Pretraining-Corpus

1 models list it as training data.

Open diffutron/DiffutronLM-Pretraining-Corpus on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.