shijiezhou/VLM4D download history
shijiezhou/VLM4D is a video text to text dataset on the Hugging Face Hub. In the last 30 days it was downloaded 985 times (106 in the last 7 days), and 21,958 times in total. It ranks #19,549 among datasets by monthly downloads.
VLM4D VLM4D is a benchmark for evaluating the spatiotemporal reasoning capabilities of Vision Language Models (VLMs). It contains real and synthetic videos paired with multiple-choice questions that require models to reason about translation, rotation, perspective, motion continuity,