beyond/chinese_clean_passages_80m download history

beyond/chinese_clean_passages_80m is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 4,161 times (479 in the last 7 days), and 24,703 times in total. It ranks #6,578 among datasets by monthly downloads.

chinese_clean_passages_80m 包含8千余万(88328203)个纯净中文段落,不包含任何字母、数字。Containing more than 80 million pure & clean Chinese passages, without any letters/digits/special tokens. 文本长度大部分介于50~200个汉字之间。The passage length is approximately 50~200 Chinese characters. 通过datasets.load_dataset()下载数据,会产生38个大

Models trained on chinese_clean_passages_80m

4 models list it as training data.

Open beyond/chinese_clean_passages_80m on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.