hasankursun/bulgarian-corpus-33b download history

hasankursun/bulgarian-corpus-33b is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 739 times (124 in the last 7 days), and 5,657 times in total. It ranks #24,342 among datasets by monthly downloads.

Bulgarian Corpus 33B (BC-33B) Dataset Summary The Bulgarian Corpus 33B (BC-33B) is a massive-scale, deduplicated, and cleaned dataset designed for training Foundation Models in Bulgarian. Comprising approximately 33.4 Billion tokens (measured with Qwen 2.5/Llama-3 token

Models trained on bulgarian-corpus-33b

2 models list it as training data.

Open hasankursun/bulgarian-corpus-33b on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.