dignity045/Collective-Corpus download history
dignity045/Collective-Corpus is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 744 times (146 in the last 7 days), and 5,263 times in total. It ranks #24,223 among datasets by monthly downloads.
🧠 Collective Corpus — Universal Pretraining + Finetuning Dataset (500B+ Tokens) Collective-Corpus is a massive-scale, multi-domain dataset designed to train Transformer-based language models from scratch and finetune them across a wide variety of domains — all in one place. 📚
Open dignity045/Collective-Corpus on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.