NLPC-UOM/document_alignment_dataset-Sinhala-Tamil-English download history

NLPC-UOM/document_alignment_dataset-Sinhala-Tamil-English is a sentence similarity dataset on the Hugging Face Hub. In the last 30 days it was downloaded 440 times (319 in the last 7 days), and 13,553 times in total. It ranks #36,763 among datasets by monthly downloads.

Dataset summary This is a gold-standard benchmark dataset for document alignment, between Sinhala-English-Tamil languages. Data had been crawled from the following news websites. News Source url Army https://www.army.lk/ Hiru http://www.hirunews.lk ITN https://www.newsf

Models trained on document_alignment_dataset-Sinhala-Tamil-English

1 models list it as training data.

Open NLPC-UOM/document_alignment_dataset-Sinhala-Tamil-English on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.