malaysia-ai/pretrain-text-dataset download history

malaysia-ai/pretrain-text-dataset is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 1,680 times (859 in the last 7 days), and 27,140 times in total. It ranks #13,026 among datasets by monthly downloads.

Dataset Introduction This dataset is a collection of malaysian texts in the Malay, English, Chinese, and Tamil languages, gathered by Malaysia AI volunteers through web crawling of malaysian websites. The dataset amounts to approximately 250 GB of text data, and has undergone deduplicati

Open malaysia-ai/pretrain-text-dataset on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.