ai4bharat/samanantar download history

ai4bharat/samanantar is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 2,416 times (768 in the last 7 days), and 93,566 times in total. It ranks #9,792 among datasets by monthly downloads.

Dataset Card for Samanantar Dataset Summary Samanantar is the largest publicly available parallel corpora collection for Indic language: Assamese, Bengali, Gujarati, Hindi, Kannada, Malayalam, Marathi, Oriya, Punjabi, Tamil, Telugu. The corpus has 49.6M sentence pairs between E

Models trained on samanantar

36 models list it as training data.

Open ai4bharat/samanantar on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.