catherinearnett/montok download history

catherinearnett/montok is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 3,435 times (1,463 in the last 7 days), and 221,182 times in total. It ranks #7,539 among datasets by monthly downloads.

MonTok: A Suite of Monolingual Tokenizers This is a set of monolingual tokenizers for 98 languages. For each language, there are Unigram, BPE, and SuperBPE tokenizers, ranging in vocabulary size from around 6k to over 200k. Training Details Training Data All tokenize

Open catherinearnett/montok on Hugging Face