PaDT-MLLM/ReferringImageCaptioning download history

PaDT-MLLM/ReferringImageCaptioning is an image to text dataset on the Hugging Face Hub. In the last 30 days it was downloaded 90 times (14 in the last 7 days), and 3,153 times in total. It ranks #122,709 among datasets by monthly downloads.

Patch-as-Decodable-Token: Towards Unified Multi-Modal Vision Tasks in MLLMs [๐Ÿ”— Released Code] [๐Ÿค— Datasets] [๐Ÿค— Checkpoints] [๐Ÿ“„ Tech Report] [๐Ÿค— Paper] Figure A. PaDT pipeline. ๐ŸŒŸ Introduction We are pleased to introduce Patch-as-Decodable Token (PaDT), a unified paradigm that enables

Open PaDT-MLLM/ReferringImageCaptioning on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.