NLPC-UOM/sentence_alignment_dataset-Sinhala-Tamil-English download history
NLPC-UOM/sentence_alignment_dataset-Sinhala-Tamil-English is a sentence similarity dataset on the Hugging Face Hub. In the last 30 days it was downloaded 589 times (125 in the last 7 days), and 56,059 times in total. It ranks #29,127 among datasets by monthly downloads.
Dataset summary This is a gold-standard benchmark dataset for sentence alignment, between Sinhala-English-Tamil languages. Data had been crawled from the following news websites. The aligned documents annotated in the dataset NLPC-UOM/document_alignment_dataset-Sinhala-Tamil-English had b
Open NLPC-UOM/sentence_alignment_dataset-Sinhala-Tamil-English on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.