repleeka/tagin-monolingual-corpus download history

repleeka/tagin-monolingual-corpus is a fill mask dataset on the Hugging Face Hub. In the last 30 days it was downloaded 22 times (2 in the last 7 days), and 22 times in total. It ranks #349,636 among datasets by monthly downloads.

Tagin Monolingual Corpus A ~110,000-segment monolingual text pool for Tagin, an endangered, highly agglutinative Tani language spoken in Arunachal Pradesh, India. The corpus was built to support Masked Language Modeling (MLM) / Continual Pre-Training (Domain-Adaptive Pre-Training, DAP

Open repleeka/tagin-monolingual-corpus on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.