vialibre/splittedspanish3bwc download history
vialibre/splittedspanish3bwc is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 185 times (61 in the last 7 days), and 77,428 times in total. It ranks #71,829 among datasets by monthly downloads.
Dataset Card for Unannotated Spanish 3 Billion Words Corpora Dataset Summary Number of lines: 300904000 (300M) Number of tokens: 2996016962 (3B) Number of chars: 18431160978 (18.4B) Languages Spanish Source Data Available to download here: Zenodo
Open vialibre/splittedspanish3bwc on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.