rijuludar/slm-tokenizer-32k download history

rijuludar/slm-tokenizer-32k is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 44 times (8 in the last 7 days), and 258 times in total. It ranks #203,866 among datasets by monthly downloads.

Custom 32k SLM Tokenizer This dataset repository contains a custom 32,768-token BPE tokenizer trained for Small Language Models (SLMs). It was created using the regex pre-tokenizer rules from Qwen/Qwen2.5-Coder-0.5B. Tokenizer Details Total Vocabulary Size: 32,768 BPE

Open rijuludar/slm-tokenizer-32k on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.