statmt/cc100 download history

statmt/cc100 is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 1,462 times (285 in the last 7 days), and 311,988 times in total. It ranks #14,534 among datasets by monthly downloads.

This corpus is an attempt to recreate the dataset used for training XLM-R. This corpus comprises of monolingual data for 100+ languages and also includes data for romanized languages (indicated by *_rom). This was constructed using the urls and paragraph indices provided by the CC-Net repository by

Models trained on cc100

55 models list it as training data.

Open statmt/cc100 on Hugging Face