oscar-corpus/oscar download history
oscar-corpus/oscar is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 632 times (146 in the last 7 days), and 2,048,734 times in total. It ranks #27,464 among datasets by monthly downloads.
The Open Super-large Crawled ALMAnaCH coRpus is a huge multilingual corpus obtained by language classification and filtering of the Common Crawl corpus using the goclassy architecture.\
Models trained on oscar
13 models list it as training data.
- LennartKeller/longformer-gottbert-base-8192-aw512 78 downloads in 30 days
- rasyosef/Llama-3.2-400M-Amharic 68 downloads in 30 days
- PantagrueLLM/text-base-camtok-oscar 44 downloads in 30 days
- PantagrueLLM/Text_Base_FR_OSCAR 40 downloads in 30 days
- salakash/SamKash-Tolstoy 25 downloads in 30 days
- sofia-uni/toxic-bert-bg 13 downloads in 30 days
- TUKE-KEMT/slovak-t5-base 13 downloads in 30 days
- abhi11nav/sakhi-telugu-681M-pretrained-0625 10 downloads in 30 days
- unbias/PasDePitiePourLeCroissantLLMBaseTest 0 downloads in 30 days
- Mika02/mistral-7b-italian-chat – downloads in 30 days
- szkelo/ggg – downloads in 30 days
- OrchestraGPT2/NekitAI – downloads in 30 days