MaLA-LM/mala-monolingual-filter download history

MaLA-LM/mala-monolingual-filter is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 4,071 times (98 in the last 7 days), and 71,265 times in total. It ranks #6,663 among datasets by monthly downloads.

MaLA Corpus: Massive Language Adaptation Corpus This is a cleaned version with some necessary data cleaning. Dataset Summary The MaLA Corpus (Massive Language Adaptation) is a comprehensive, multilingual dataset designed to support the continual pre-training of large l

Open MaLA-LM/mala-monolingual-filter on Hugging Face