eduagarcia/cc100-pt download history

eduagarcia/cc100-pt is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 449 times (88 in the last 7 days), and 1,155 times in total. It ranks #36,061 among datasets by monthly downloads.

C100-PT CC100-PT is the is the portuguese subset from C100. C100 was created for training the multilingual Transformer XLM-R, containing two terabytes of cleaned data from 2018 snapshots of the Common Crawl project in 100 languages.

Open eduagarcia/cc100-pt on Hugging Face