PredictiveManish/multilingual-corpus download history
PredictiveManish/multilingual-corpus is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 80 times (16 in the last 7 days), and 502 times in total. It ranks #134,522 among datasets by monthly downloads.
Filtered Multilingual Corpus (EN-HI-PA) Dataset Description A cleaned and balanced multilingual corpus extracted from the Samanantar dataset, containing parallel sentences in English, Hindi, and Punjabi. Languages The dataset contains text in three languages: Englis
Models trained on multilingual-corpus
1 models list it as training data.
- PredictiveManish/Trimurti-LM 20 downloads in 30 days
Open PredictiveManish/multilingual-corpus on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.