nltk-data-hub/crubadan download history

nltk-data-hub/crubadan is a text classification dataset on the Hugging Face Hub. In the last 30 days it was downloaded 88 times (5 in the last 7 days), and 1,596 times in total. It ranks #124,722 among datasets by monthly downloads.

NLTK Crúbadán Language ID Corpus Character 3-gram frequency tables for 449 writing systems, collected by Kevin Scannell's An Crúbadán web crawler (2010). Distributed via NLTK. Trigrams use < (word start) and > (word end) as boundary markers. Configs Config Description Sch

Open nltk-data-hub/crubadan on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.