GPT-NL/GPT-NL_Public_Corpus download history
GPT-NL/GPT-NL_Public_Corpus is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 2,954 times (819 in the last 7 days), and 49,048 times in total. It ranks #8,462 among datasets by monthly downloads.
Dataset Card GPT-NL Public Corpus The GPT-NL Public Corpus is the largest permissively licensed Dutch-language resource available for large language model pretraining. It consists of 29 curated collections totaling over 524 billion tokens, including 36B Dutch, 207B English, 232B code,
Models trained on GPT-NL_Public_Corpus
4 models list it as training data.
- mradermacher/walnoot-8b-instruct-i1-GGUF 2.9K downloads in 30 days
- WAINUT/walnoot-8b-instruct 1.5K downloads in 30 days
- mradermacher/walnoot-8b-instruct-GGUF 913 downloads in 30 days
- kamoo-ai/kamoo-one-135m 148 downloads in 30 days
Open GPT-NL/GPT-NL_Public_Corpus on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.