atahanuz/setimes-en-tr-aligned-corpus download history

atahanuz/setimes-en-tr-aligned-corpus is a translation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 96 times (40 in the last 7 days), and 251 times in total. It ranks #117,245 among datasets by monthly downloads.

SETimes EN-TR — Sentence-Aligned, LLM-Cleaned A cleaned and re-aligned version of the SETimes English-Turkish parallel corpus. The original SETimes data is paragraph-style — each "pair" can contain a headline, a dateline, several body sentences, and a source citation, all glued togeth

Open atahanuz/setimes-en-tr-aligned-corpus on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.