nikolina-p/gutenberg_clean_tokenized_en_splits download history

nikolina-p/gutenberg_clean_tokenized_en_splits is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 7 times (3 in the last 7 days), and 957 times in total. It ranks #677,853 among datasets by monthly downloads.

Overview This dataset is a tokenized version of the cleaned English-language subset of the Project Gutenberg Dataset manu/project_gutenberg. It contains full-text books in English, free of boilerplate content and duplicates, and includes a pre-tokenized version of each book's content usin

Open nikolina-p/gutenberg_clean_tokenized_en_splits on Hugging Face