ThingAI/Italian-Common-Corpus download history
ThingAI/Italian-Common-Corpus is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 69 times (12 in the last 7 days), and 224 times in total. It ranks #149,391 among datasets by monthly downloads.
Italian-Common-Corpus The Italian dataset with the highest density of useful information per token. Built by ModotAI for training Italian language models. Subsets Subset File Documents Words Description Web Crawl icc-web.parquet ~27K ~17M Italian sources: new
Open ThingAI/Italian-Common-Corpus on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.