thng292/fineweb-subset-10M download history
thng292/fineweb-subset-10M is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 275 times (93 in the last 7 days), and 3,159 times in total. It ranks #52,566 among datasets by monthly downloads.
This dataset comes from HuggingFaceFW/fineweb-2 and HuggingFaceFW/fineweb-edu. It includes five languages: Vietnamese, English, French, Japanese, and Chinese. Each language has 2,000,000 training samples and 100,000 test samples, totaling 10 million rows for training and 500,000 rows for testing. Th
Open thng292/fineweb-subset-10M on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.