bertin-project/mc4-es-sampled download history

bertin-project/mc4-es-sampled is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 1,315 times (692 in the last 7 days), and 37,559 times in total. It ranks #15,757 among datasets by monthly downloads.

50 million documents in Spanish extracted from mC4 applying perplexity sampling via mc4-sampling: "https://huggingface.co/datasets/bertin-project/mc4-sampling". Please, refer to BERTIN Project. The original dataset is the Multlingual Colossal, Cleaned version of Common Crawl's web crawl corpus (mC4)

Models trained on mc4-es-sampled

6 models list it as training data.

Open bertin-project/mc4-es-sampled on Hugging Face