chauhan45/IndicCorpV2 download history
chauhan45/IndicCorpV2 is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 869 times (6 in the last 7 days), and 1,038 times in total. It ranks #21,470 among datasets by monthly downloads.
IndicCorp v2 Dataset Towards Leaving No Indic Language Behind: Building Monolingual Corpora, Benchmark and Models for Indic Languages This repository contains the pretraining data for the paper published at ACL 2023. Example Usage from datasets import load_dataset