ChamaraVishwajithRajapaksha/sinhala-text-dataset download history

ChamaraVishwajithRajapaksha/sinhala-text-dataset is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 27 times (4 in the last 7 days), and 114 times in total. It ranks #297,781 among datasets by monthly downloads.

Sinhala Continuous Pretraining Corpus Curated by HelaAI Dataset Summary This dataset is a Sinhala-language text corpus assembled for continuous pretraining of language models. It combines multiple sources into a single, cleaned, block-structured corpus: News articles —

Models trained on sinhala-text-dataset

1 models list it as training data.

Open ChamaraVishwajithRajapaksha/sinhala-text-dataset on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.