GPT-NL/GPT-NL_Public_Corpus download history

GPT-NL/GPT-NL_Public_Corpus is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 2,954 times (819 in the last 7 days), and 49,048 times in total. It ranks #8,462 among datasets by monthly downloads.

Dataset Card GPT-NL Public Corpus The GPT-NL Public Corpus is the largest permissively licensed Dutch-language resource available for large language model pretraining. It consists of 29 curated collections totaling over 524 billion tokens, including 36B Dutch, 207B English, 232B code,

Models trained on GPT-NL_Public_Corpus

4 models list it as training data.

Open GPT-NL/GPT-NL_Public_Corpus on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.