gsarti/clean_mc4_it download history

gsarti/clean_mc4_it is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 2,115 times (383 in the last 7 days), and 50,345 times in total. It ranks #10,812 among datasets by monthly downloads.

A thoroughly cleaned version of the Italian portion of the multilingual colossal, cleaned version of Common Crawl's web crawl corpus (mC4) by AllenAI. Based on Common Crawl dataset: "https://commoncrawl.org". This is the processed version of Google's mC4 dataset by AllenAI, with further cleaning

Models trained on clean_mc4_it

26 models list it as training data.

Open gsarti/clean_mc4_it on Hugging Face