llamaindex/vdr-multilingual-train download history

llamaindex/vdr-multilingual-train is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 1,406 times (547 in the last 7 days), and 47,828 times in total. It ranks #14,949 among datasets by monthly downloads.

Multilingual Visual Document Retrieval Dataset This dataset consists of 500k multilingual query image samples, collected and generated from scratch using public internet pdfs. The queries are synthetic and generated using VLMs (gemini-1.5-pro and Qwen2-VL-72B). It was used to train the

Models trained on vdr-multilingual-train

45 models list it as training data.

Open llamaindex/vdr-multilingual-train on Hugging Face