oscar-corpus/OSCAR-2109 download history

oscar-corpus/OSCAR-2109 is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 83 times (16 in the last 7 days), and 405,523 times in total. It ranks #131,075 among datasets by monthly downloads.

The Open Super-large Crawled Aggregated coRpus is a huge multilingual corpus obtained by language classification and filtering of the Common Crawl corpus using the goclassy architecture.\

Models trained on OSCAR-2109

574 models list it as training data.

Spaces using OSCAR-2109

Open oscar-corpus/OSCAR-2109 on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.