MaLA-LM/mala-monolingual-integration download history

MaLA-LM/mala-monolingual-integration is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 4,064 times (596 in the last 7 days), and 30,418 times in total. It ranks #6,668 among datasets by monthly downloads.

MaLA Corpus: Massive Language Adaptation Corpus This is the noisy version that integrates texts from different sources. Dataset Summary The MaLA Corpus (Massive Language Adaptation) is a comprehensive, multilingual dataset designed to support the continual pre-training

Open MaLA-LM/mala-monolingual-integration on Hugging Face