uctnlp/mzansi-text-tokenized download history
uctnlp/mzansi-text-tokenized is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 393 times (132 in the last 7 days), and 4,484 times in total. It ranks #39,920 among datasets by monthly downloads.
MzansiText Tokenized Document-tokenized MzansiText used by the later SALLM pretraining runs. Tokenizer: custom 65,536-vocabulary BPE tokenizer Splits: 3,943,584 train, 19,379 validation, and 19,341 test rows Each raw document remains one row and is truncated to at most 2,048 tokens S