thng292/fineweb-subset-1M download history

thng292/fineweb-subset-1M is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 486 times (351 in the last 7 days), and 2,193 times in total. It ranks #33,888 among datasets by monthly downloads.

This dataset comes from HuggingFaceFW/fineweb-2 and HuggingFaceFW/fineweb-edu. It includes five languages: Vietnamese, English, French, Japanese, and Chinese. Each language has 200,000 training samples and 10,000 test samples, totaling 1 million rows for training and 50,000 rows for testing. The def

Models trained on fineweb-subset-1M

1 models list it as training data.

Open thng292/fineweb-subset-1M on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.