SEACrowd/cc100 download history

SEACrowd/cc100 is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 124 times (58 in the last 7 days), and 2,226 times in total. It ranks #97,463 among datasets by monthly downloads.

This corpus is an attempt to recreate the dataset used for training XLM-R. This corpus comprises of monolingual data for 100+ languages and also includes data for romanized languages (indicated by *_rom). This was constructed using the urls and paragraph indices provi

Open SEACrowd/cc100 on Hugging Face