crscardellino/spanish_billion_words download history

crscardellino/spanish_billion_words is an other dataset on the Hugging Face Hub. In the last 30 days it was downloaded 180 times (52 in the last 7 days), and 15,837 times in total. It ranks #73,655 among datasets by monthly downloads.

An unannotated Spanish corpus of nearly 1.5 billion words, compiled from different resources from the web. This resources include the spanish portions of SenSem, the Ancora Corpus, some OPUS Project Corpora and the Europarl, the Tibidabo Treebank, the IULA Spanish LSP Treebank, and dumps from the Sp

Open crscardellino/spanish_billion_words on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.