ArmelR/the-pile-splitted download history

ArmelR/the-pile-splitted is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 18,366 times (4,788 in the last 7 days), and 231,850 times in total. It ranks #1,695 among datasets by monthly downloads.

Dataset description The pile is an 800GB dataset of english text designed by EleutherAI to train large-scale language models. The original version of the dataset can be found here. The dataset is divided into 22 smaller high-quality datasets. For more information each of them, please re

Models trained on the-pile-splitted

1 models list it as training data.

Spaces using the-pile-splitted

Open ArmelR/the-pile-splitted on Hugging Face