deepcs233/Visual-CoT download history

deepcs233/Visual-CoT is an image text to text dataset on the Hugging Face Hub. In the last 30 days it was downloaded 7,363 times (1,164 in the last 7 days), and 64,706 times in total. It ranks #4,056 among datasets by monthly downloads.

VisCoT Dataset Card There is a shortage of multimodal datasets for training multi-modal large language models (MLLMs) that require to identify specific regions in an image for additional attention to improve response performance. This type of dataset with grounding bbox annotations coul

Spaces using Visual-CoT

Open deepcs233/Visual-CoT on Hugging Face