LVSTCK/macedonian-corpus-cleaned download history
LVSTCK/macedonian-corpus-cleaned is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 47 times (8 in the last 7 days), and 1,404 times in total. It ranks #194,537 among datasets by monthly downloads.
Macedonian Corpus - Cleaned raw version here Paper 🌟 Key Highlights Size: 35.5 GB, Word Count: 3.31 billion Filtered for irrelevant and low-quality content using C4 and Gopher filtering. Includes text from 10+ sources such as fineweb-2, HPLT-2, Wikipedia, and more. 📋
Open LVSTCK/macedonian-corpus-cleaned on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.