oahegiaerhg/common_corpus download history

oahegiaerhg/common_corpus is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 3,825 times (505 in the last 7 days), and 5,368 times in total. It ranks #6,961 among datasets by monthly downloads.

Common Corpus Full paper - ICLR 2026 oral Common Corpus is the largest open licensed text dataset, comprising 2.27 trillion tokens (2,267,302,720,836 tokens). It is a diverse dataset, consisting of books, newspapers, scientific articles, government and legal documents, code, and

Open oahegiaerhg/common_corpus on Hugging Face