NLPC-UOM/document_alignment_dataset-Sinhala-Tamil-English download history
NLPC-UOM/document_alignment_dataset-Sinhala-Tamil-English is a sentence similarity dataset on the Hugging Face Hub. In the last 30 days it was downloaded 440 times (319 in the last 7 days), and 13,553 times in total. It ranks #36,763 among datasets by monthly downloads.
Dataset summary This is a gold-standard benchmark dataset for document alignment, between Sinhala-English-Tamil languages. Data had been crawled from the following news websites. News Source url Army https://www.army.lk/ Hiru http://www.hirunews.lk ITN https://www.newsf
Models trained on document_alignment_dataset-Sinhala-Tamil-English
1 models list it as training data.
- yeezer/gpt2-tamil-merged 16 downloads in 30 days
Open NLPC-UOM/document_alignment_dataset-Sinhala-Tamil-English on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.