cambridge-climb/BabyLM download history

cambridge-climb/BabyLM is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 3,002 times (581 in the last 7 days), and 37,881 times in total. It ranks #8,332 among datasets by monthly downloads.

Dataset for the shared baby language modeling task. The goal is to train a language model from scratch on this data which represents roughly the amount of text and speech data a young child observes.

Models trained on BabyLM

3 models list it as training data.

Open cambridge-climb/BabyLM on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.