hasankursun/dutch-corpus-200b download history

hasankursun/dutch-corpus-200b is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 492 times (132 in the last 7 days), and 5,151 times in total. It ranks #33,549 among datasets by monthly downloads.

Dutch Corpus 200B (DC-200B) Dataset Summary The Dutch Corpus 200B (DC-200B) is the largest open-source, deduplicated, and professionally cleaned dataset designed for training Foundation Models in the Dutch language. Comprising approximately 202 Billion tokens (measured

Open hasankursun/dutch-corpus-200b on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.