sinhala-nlp/sinhala-1.5B-corpus download history
sinhala-nlp/sinhala-1.5B-corpus is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 66 times (30 in the last 7 days), and 1,357 times in total. It ranks #152,630 among datasets by monthly downloads.
Sinhala-1.5B-corpus is a large-scale monolingual corpus for Sinhala, released as part of the ACL 2025 paper "Sinhala Encoder-only Language Models and Evaluation". The corpus was created to address the limited availability of large-scale pre-training resources for Sinhala. It combines Sinhala text fr
Open sinhala-nlp/sinhala-1.5B-corpus on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.