ai4bharat/IndicCorpV2 download history
ai4bharat/IndicCorpV2 is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 3,592 times (1,195 in the last 7 days), and 35,476 times in total. It ranks #7,297 among datasets by monthly downloads.
IndicCorp v2 Dataset Towards Leaving No Indic Language Behind: Building Monolingual Corpora, Benchmark and Models for Indic Languages This repository contains the pretraining data for the paper published at ACL 2023. Example Usage from datasets import l
Models trained on IndicCorpV2
12 models list it as training data.
- ai4bharat/Cadence-Fast 6.3K downloads in 30 days
- ai4bharat/Cadence 3.6K downloads in 30 days
- ai4bharat/IndicBERT-v3-270M 618 downloads in 30 days
- sraivante/Sanskrit-322M-Base 459 downloads in 30 days
- samvaran/chandohasam 284 downloads in 30 days
- ai4bharat/IndicBERT-v3-4B 182 downloads in 30 days
- ai4bharat/IndicBERT-v3-1B 169 downloads in 30 days
- kkkamur07/hindi-modernbert 46 downloads in 30 days
- BERTCHEESIE/Kannada-tokenizer 23 downloads in 30 days
- ishathombre/monolingual-hindi-from-scratch 18 downloads in 30 days
- BERTCHEESIE/KannaBERT-xl 7 downloads in 30 days
- abirmaheshwari/abirmarv1 – downloads in 30 days
Open ai4bharat/IndicCorpV2 on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.