eduagarcia/multilingual_tokenizer_benchmark download history
eduagarcia/multilingual_tokenizer_benchmark is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 120 times (24 in the last 7 days), and 731 times in total. It ranks #99,950 among datasets by monthly downloads.
Multilingual Tokenizer Benchmark More details of each subset like word count, character count, original sources, etc, can be found in the dataset_meta.yaml file in the repository root. Natural language word count functions Download spacy models pip install ntlk spacy pygments u
Models trained on multilingual_tokenizer_benchmark
1 models list it as training data.
- dawncr0w/KorByte-128K – downloads in 30 days
Open eduagarcia/multilingual_tokenizer_benchmark on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.