datablations/oscar-subsets download history
datablations/oscar-subsets is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 482 times (115 in the last 7 days), and 20,559 times in total. It ranks #34,104 among datasets by monthly downloads.
Dataset Summary Various subsets of the English OSCAR with different numbers of tokens measured with the GPT2Tokenizer. This data is used in the paper Scaling Data-Constrained Language Models. Please refer to our GitHub repository for more details. @article{muennighoff2023scaling, title