chinese-babylm-org/babylm-zho-100M download history

chinese-babylm-org/babylm-zho-100M is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 68 times (10 in the last 7 days), and 1,519 times in total. It ranks #150,988 among datasets by monthly downloads.

babylm-zho-100M A filtered version of BabyLM-community/babylm-zho, a Chinese-language corpus designed for the BabyLM Challenge. This is the official training data for Chinese BabyLM Challenge. Size The filtered dataset contains approximately 101,343,320 tokens (tokenize

Models trained on babylm-zho-100M

1 models list it as training data.

Open chinese-babylm-org/babylm-zho-100M on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.