wizzense/Usenet-Corpus-1980-2013 download history
wizzense/Usenet-Corpus-1980-2013 is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 61 times (14 in the last 7 days), and 228 times in total. It ranks #162,562 among datasets by monthly downloads.
Usenet Corpus 1980–2013 One of the largest curated Usenet archives available for AI training — 103.1 billion tokens of authentic pre-LLM human discourse. This dataset contains 408 million cleaned and deduplicated Usenet posts spanning 1980–2013 across 18,347 newsgroups. It is sourced from
Open wizzense/Usenet-Corpus-1980-2013 on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.