procesaur/sr-tokenizer-test download history
procesaur/sr-tokenizer-test is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 107 times (38 in the last 7 days), and 960 times in total. It ranks #108,408 among datasets by monthly downloads.
Sr Tokenizer test This dataset provides a large Serbian text corpus designed for training and evaluating of tokenizers for Serbian language models. It combines multiple sources of Serbian text in both Cyrillic and Latin scripts, unified into a consistent JSONL format with id and text fie
Models trained on sr-tokenizer-test
1 models list it as training data.
- procesaur/Srna_tokenizer 17 downloads in 30 days
Open procesaur/sr-tokenizer-test on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.