Shitao/MLDR download history

Shitao/MLDR is a text retrieval dataset on the Hugging Face Hub. In the last 30 days it was downloaded 1,730 times (344 in the last 7 days), and 253,922 times in total. It ranks #12,694 among datasets by monthly downloads.

Dataset Summary MLDR is a Multilingual Long-Document Retrieval dataset built on Wikipeida, Wudao and mC4, covering 13 typologically diverse languages. Specifically, we sample lengthy articles from Wikipedia, Wudao and mC4 datasets and randomly choose paragraphs from them. Then we use GPT-

Models trained on MLDR

1 models list it as training data.

Open Shitao/MLDR on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.