text-machine-lab/vocab_filtered_dataset_2.1B download history
text-machine-lab/vocab_filtered_dataset_2.1B is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 265 times (18 in the last 7 days), and 2,712 times in total. It ranks #54,094 among datasets by monthly downloads.
Dataset Card for "vocab_filtered_dataset_2.1B" Dataset Summary This data is the simplified vocabulary-filtered pretraining data published by "Emergent Abilities in Reduced-Scale Generative Language Models". The vocabulary is derived from the AO-Childes speech corpus (https://gi
Open text-machine-lab/vocab_filtered_dataset_2.1B on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.