NNEngine/Gutenberg-Clean-40M download history
NNEngine/Gutenberg-Clean-40M is a text classification dataset on the Hugging Face Hub. In the last 30 days it was downloaded 30 times (12 in the last 7 days), and 378 times in total. It ranks #274,414 among datasets by monthly downloads.
📚 TinyWay-Gutenberg-Clean-40M A large-scale, high-quality English text dataset derived from Project Gutenberg, cleaned, normalized, deduplicated, and segmented into fixed-length samples for efficient language model pretraining. This dataset is designed to support training small and medium
Open NNEngine/Gutenberg-Clean-40M on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.