raphaelmerx/MADLAD-400-Tetun download history

raphaelmerx/MADLAD-400-Tetun is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 23 times (6 in the last 7 days), and 572 times in total. It ranks #338,096 among datasets by monthly downloads.

MADLAD-400-Tetun The Tetun Dili "clean" split of the MADLAD-400 dataset, sentencized and translated to English using MADLAD-400 3b. Each row has: text: the original text from MADLAD sentences: the text sentencized to Tetun, using Moses, which has Tetun non-breaking prefixes. sentences_e

Open raphaelmerx/MADLAD-400-Tetun on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.