jhu-clsp/paradocs download history
jhu-clsp/paradocs is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 579 times (31 in the last 7 days), and 12,910 times in total. It ranks #29,464 among datasets by monthly downloads.
ParaDocs is a multilingual machine translation dataset that has labelled document annotations for ParaCrawl, NewsCommentary, and Europarl data which can be used to create parallel document datasets for training of context-aware machine translation models.