projecte-aina/CATalog download history

projecte-aina/CATalog is a fill mask dataset on the Hugging Face Hub. In the last 30 days it was downloaded 3,178 times (538 in the last 7 days), and 45,951 times in total. It ranks #7,988 among datasets by monthly downloads.

Dataset Summary CATalog is a diverse, open-source Catalan corpus for language modelling. It consists of text documents from 26 different sources, including web crawling, news, forums, digital libraries and public institutions, totaling in 17.45 billion words. Supported Tasks and L

Models trained on CATalog

33 models list it as training data.

Open projecte-aina/CATalog on Hugging Face