PleIAs/common_corpus download history
PleIAs/common_corpus is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 175,700 times (26,539 in the last 7 days), and 1,920,733 times in total. It ranks #135 among datasets by monthly downloads.
Common Corpus Full paper - ICLR 2026 oral Common Corpus is the largest open licensed text dataset, comprising 2.27 trillion tokens (2,267,302,720,836 tokens). It is a diverse dataset, consisting of books, newspapers, scientific articles, government and legal documents, code, and more
Models trained on common_corpus
38 models list it as training data.
- PleIAs/Pleias-1.2b-Preview 3.3K downloads in 30 days
- PleIAs/Pleias-350m-Preview 3.1K downloads in 30 days
- PleIAs/Pleias-3b-Preview 3.1K downloads in 30 days
- mradermacher/Unbound-v1.12.0-27B-i1-GGUF 1.4K downloads in 30 days
- mradermacher/Mira-v1.12.1-27B-i1-GGUF 1.2K downloads in 30 days
- mittagessen/bytellama-43m-cc 1.1K downloads in 30 days
- mradermacher/rodin-1b-instruct-i1-GGUF 667 downloads in 30 days
- mradermacher/Unbound-v1.12.0-27B-GGUF 412 downloads in 30 days
- mradermacher/rodin-1b-i1-GGUF 407 downloads in 30 days
- mradermacher/Mira-v1.12.1-27B-GGUF 393 downloads in 30 days
- rodin-llm/rodin-1b 261 downloads in 30 days
- rodin-llm/rodin-1b-instruct 220 downloads in 30 days
Spaces using common_corpus
- Model Pulse 8 likes
Open PleIAs/common_corpus on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.