statmt/cc100 download history
statmt/cc100 is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 1,462 times (285 in the last 7 days), and 311,988 times in total. It ranks #14,534 among datasets by monthly downloads.
This corpus is an attempt to recreate the dataset used for training XLM-R. This corpus comprises of monolingual data for 100+ languages and also includes data for romanized languages (indicated by *_rom). This was constructed using the urls and paragraph indices provided by the CC-Net repository by
Models trained on cc100
55 models list it as training data.
- nomic-ai/nomic-xlm-2048 475 downloads in 30 days
- dsfsi/BantuBERTa 189 downloads in 30 days
- alexander-sh/mDeBERTa-v3-multi-sent 125 downloads in 30 days
- KoichiYasuoka/modernbert-base-thai-cc100 107 downloads in 30 days
- goldfish-models/lim_latn_100mb 27 downloads in 30 days
- goldfish-models/lug_latn_5mb 26 downloads in 30 days
- goldfish-models/ssw_latn_5mb 25 downloads in 30 days
- goldfish-models/que_latn_full 25 downloads in 30 days
- goldfish-models/nso_latn_full 24 downloads in 30 days
- goldfish-models/que_latn_10mb 24 downloads in 30 days
- goldfish-models/lin_latn_full 23 downloads in 30 days
- goldfish-models/lug_latn_10mb 23 downloads in 30 days