helloadhavan/CC-FilteredCorpus download history

helloadhavan/CC-FilteredCorpus is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 170 times (23 in the last 7 days), and 1,054 times in total. It ranks #76,958 among datasets by monthly downloads.

English Cleaned Common Crawl Markdown Dataset An English-focused dataset created from Common Crawl, cleaned and converted to Markdown. The goal is to preserve web-document structure so that AI models can learn both natural language and Markdown formatting. Features Eng

Open helloadhavan/CC-FilteredCorpus on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.