projecte-aina/CATalog download history
projecte-aina/CATalog is a fill mask dataset on the Hugging Face Hub. In the last 30 days it was downloaded 3,178 times (538 in the last 7 days), and 45,951 times in total. It ranks #7,988 among datasets by monthly downloads.
Dataset Summary CATalog is a diverse, open-source Catalan corpus for language modelling. It consists of text documents from 26 different sources, including web crawling, news, forums, digital libraries and public institutions, totaling in 17.45 billion words. Supported Tasks and L
Models trained on CATalog
33 models list it as training data.
- BSC-LT/salamandra-7b-instruct 26.4K downloads in 30 days
- BSC-LT/salamandra-2b-instruct 3.2K downloads in 30 days
- BSC-LT/ALIA-40b 1.8K downloads in 30 days
- BSC-LT/salamandra-7b 1.7K downloads in 30 days
- BSC-LT/salamandra-2b 1.6K downloads in 30 days
- mradermacher/ALIA-40b-i1-GGUF 1.2K downloads in 30 days
- mradermacher/salamandra-2b-instruct-i1-GGUF 879 downloads in 30 days
- mradermacher/salamandra-2b-instruct-GGUF 662 downloads in 30 days
- mradermacher/ALIA-40b-GGUF 492 downloads in 30 days
- tensorblock/ALIA-40b-GGUF 269 downloads in 30 days
- tensorblock/salamandra-2b-instruct-GGUF 222 downloads in 30 days
- tensorblock/salamandra-7b-instruct-GGUF 189 downloads in 30 days