oscar-corpus/OSCAR-2109 download history
oscar-corpus/OSCAR-2109 is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 83 times (16 in the last 7 days), and 405,523 times in total. It ranks #131,075 among datasets by monthly downloads.
The Open Super-large Crawled Aggregated coRpus is a huge multilingual corpus obtained by language classification and filtering of the Common Crawl corpus using the goclassy architecture.\
Models trained on OSCAR-2109
574 models list it as training data.
- goldfish-models/eng_latn_100mb 538 downloads in 30 days
- goldfish-models/spa_latn_1000mb 387 downloads in 30 days
- dragonSwing/vibert-capu 361 downloads in 30 days
- goldfish-models/eng_latn_1000mb 350 downloads in 30 days
- goldfish-models/hin_deva_1000mb 323 downloads in 30 days
- goldfish-models/urd_arab_5mb 318 downloads in 30 days
- dragonSwing/xlm-roberta-capu 316 downloads in 30 days
- vgaraujov/t5-base-spanish 249 downloads in 30 days
- goldfish-models/tha_thai_5mb 221 downloads in 30 days
- vgaraujov/bart-base-spanish 218 downloads in 30 days
- goldfish-models/gla_latn_5mb 217 downloads in 30 days
- goldfish-models/pan_guru_5mb 207 downloads in 30 days
Spaces using OSCAR-2109
- janar/toypdf 0 likes
Open oscar-corpus/OSCAR-2109 on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.