LVSTCK/macedonian-corpus-cleaned-dedup download history

LVSTCK/macedonian-corpus-cleaned-dedup is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 50 times (11 in the last 7 days), and 2,019 times in total. It ranks #186,032 among datasets by monthly downloads.

Macedonian Corpus - Cleaned and Deduplicated Paper 🌟 Key Highlights Size: 16.78 GB, Word Count: 1.47 billion Deduplicated using MinHash to remove redundant documents. 📋 Overview Macedonian is widely recognized as a low-resource language in the field of NLP. Publicl

Models trained on macedonian-corpus-cleaned-dedup

5 models list it as training data.

Open LVSTCK/macedonian-corpus-cleaned-dedup on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.