PleIAs/common_corpus download history

PleIAs/common_corpus is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 175,700 times (26,539 in the last 7 days), and 1,920,733 times in total. It ranks #135 among datasets by monthly downloads.

Common Corpus Full paper - ICLR 2026 oral Common Corpus is the largest open licensed text dataset, comprising 2.27 trillion tokens (2,267,302,720,836 tokens). It is a diverse dataset, consisting of books, newspapers, scientific articles, government and legal documents, code, and more

Models trained on common_corpus

38 models list it as training data.

Spaces using common_corpus

Open PleIAs/common_corpus on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.