MichaelR207/high-quality-cc-21b download history

MichaelR207/high-quality-cc-21b is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 617 times (346 in the last 7 days), and 2,667 times in total. It ranks #27,973 among datasets by monthly downloads.

high_quality A high-quality English web text corpus extracted from Common Crawl WARC files using an LLM-based extraction and quality pipeline. Dataset Summary high_quality is a pretraining-grade corpus of cleaned web documents. Raw Common Crawl WARC records are passed t

Open MichaelR207/high-quality-cc-21b on Hugging Face