procesaur/sr-tokenizer-test download history

procesaur/sr-tokenizer-test is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 107 times (38 in the last 7 days), and 960 times in total. It ranks #108,408 among datasets by monthly downloads.

Sr Tokenizer test This dataset provides a large Serbian text corpus designed for training and evaluating of tokenizers for Serbian language models. It combines multiple sources of Serbian text in both Cyrillic and Latin scripts, unified into a consistent JSONL format with id and text fie

Models trained on sr-tokenizer-test

1 models list it as training data.

Open procesaur/sr-tokenizer-test on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.