text-machine-lab/vocab_filtered_dataset_22B download history
text-machine-lab/vocab_filtered_dataset_22B is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 667 times (81 in the last 7 days), and 23,015 times in total. It ranks #26,353 among datasets by monthly downloads.
Dataset Card for "vocab_filtered_dataset_22B" Dataset Summary This data is the simplified vocabulary-filtered pretraining data published by "Emergent Abilities in Reduced-Scale Generative Language Models". The vocabulary is derived from the AO-Childes speech corpus (https://git
Open text-machine-lab/vocab_filtered_dataset_22B on Hugging Face