ai4bharat/IndicCorpV2 download history

ai4bharat/IndicCorpV2 is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 3,592 times (1,195 in the last 7 days), and 35,476 times in total. It ranks #7,297 among datasets by monthly downloads.

IndicCorp v2 Dataset Towards Leaving No Indic Language Behind: Building Monolingual Corpora, Benchmark and Models for Indic Languages This repository contains the pretraining data for the paper published at ACL 2023. Example Usage from datasets import l

Models trained on IndicCorpV2

12 models list it as training data.

Open ai4bharat/IndicCorpV2 on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.