bertin-project/mc4-es-sampled download history
bertin-project/mc4-es-sampled is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 1,315 times (692 in the last 7 days), and 37,559 times in total. It ranks #15,757 among datasets by monthly downloads.
50 million documents in Spanish extracted from mC4 applying perplexity sampling via mc4-sampling: "https://huggingface.co/datasets/bertin-project/mc4-sampling". Please, refer to BERTIN Project. The original dataset is the Multlingual Colossal, Cleaned version of Common Crawl's web crawl corpus (mC4)
Models trained on mc4-es-sampled
6 models list it as training data.
- bertin-project/bertin-roberta-base-spanish 2.6K downloads in 30 days
- vgaraujov/t5-base-spanish 248 downloads in 30 days
- vgaraujov/bart-base-spanish 219 downloads in 30 days
- bertin-project/bertin-gpt-j-6B 195 downloads in 30 days
- bertin-project/bertin-gpt-j-6B-infolibros 27 downloads in 30 days
- vgaraujov/led-base-16384-spanish 15 downloads in 30 days