SPAISS6F1/slm-pretrain-corpus download history

SPAISS6F1/slm-pretrain-corpus is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 26 times (14 in the last 7 days), and 197 times in total. It ranks #307,308 among datasets by monthly downloads.

SLM Pretrain Corpus (Thai) Cleaned Thai text corpus assembled for pretraining a small language model (SLM) from scratch. Built by the data_pipeline project: each upstream source is normalized to a single text field, unicode-normalized, whitespace-cleaned, length-filtered, and exact-de

Open SPAISS6F1/slm-pretrain-corpus on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.