open-index/ccrawl-urls download history

open-index/ccrawl-urls is a text retrieval dataset on the Hugging Face Hub. In the last 30 days it was downloaded 718 times (89 in the last 7 days), and 1,479 times in total. It ranks #24,870 among datasets by monthly downloads.

Common Crawl URL Index Every URL Common Crawl has seen, as a slim columnar table, ready to seed a crawler frontier What is it? This dataset is the URL-level index of Common Crawl, republished as clean Parquet. Common Crawl is a non-profit that crawls the web every mon

Open open-index/ccrawl-urls on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.