diffutron/DiffutronLM-Pretraining-Corpus download history
diffutron/DiffutronLM-Pretraining-Corpus is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 46 times (16 in the last 7 days), and 516 times in total. It ranks #197,755 among datasets by monthly downloads.
DiffutronLM-Pretraining-Corpus DiffutronLM-Pretraining-Corpus is the comprehensive, filtered Turkish text dataset used during the Continual Pre-training (CPT) phase of the Diffutron language models. The primary goal of this dataset was to align the cross-lingual representations of a mult
Models trained on DiffutronLM-Pretraining-Corpus
1 models list it as training data.
- TFLai/tifilBERT-Base 76 downloads in 30 days
Open diffutron/DiffutronLM-Pretraining-Corpus on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.