google-research-datasets/crawl_domain download history
google-research-datasets/crawl_domain is an other dataset on the Hugging Face Hub. In the last 30 days it was downloaded 183 times (41 in the last 7 days), and 14,241 times in total. It ranks #72,405 among datasets by monthly downloads.
Corpus of domain names scraped from Common Crawl and manually annotated to add word boundaries (e.g. "commoncrawl" to "common crawl"). Breaking domain names such as "openresearch" into component words "open" and "research" is important for applications such as Text-to-Speech synthesis and web search
Open google-research-datasets/crawl_domain on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.