Describe Anything Model for Visual Question Answering on Text-rich Images
cs.CV, cs.LG
Submitted: 2025-07-16
Updated: 2025-08-02
Comments: 11 pages, 5 figures. Accepted to VisionDocs @ ICCV 2025
Journal ref: Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops (ICCVW), 2025
DOI: 10.1109/ICCVW69036.2025.00776
Code: https://github.com/Linvyl/DAM-QA
License: http://creativecommons.org/licenses/by/4.0/
The gist: Recent progress has been made in region-aware vision-language modeling, particularly with the emergence of the Describe Anything Model (DAM).
Terminology
Abstract
Recent progress has been made in region-aware vision-language modeling, particularly with the emergence of the Describe Anything Model (DAM). DAM is capable of generating detailed descriptions of any specific image areas or objects without the need for additional localized image-text alignment supervision. We hypothesize that such region-level descriptive capability is beneficial for the task of Visual Question Answering (VQA), especially in challenging scenarios involving images with dense text. In such settings, the fine-grained extraction of textual information is crucial to producing correct answers. Motivated by this, we introduce DAM-QA, a framework with a tailored evaluation protocol, developed to investigate and harness the region-aware capabilities from DAM for the text-rich VQA problem that requires reasoning over text-based information within images. DAM-QA incorporates a mechanism that aggregates answers from multiple regional views of image content, enabling more effective identification of evidence that may be tied to text-related elements. Experiments on six VQA benchmarks show that our approach consistently outperforms the baseline DAM, with a notable 7+ point gain on DocVQA. DAM-QA also achieves the best overall performance among region-aware models with fewer parameters, significantly narrowing the gap with strong generalist VLMs. These results highlight the potential of DAM-like models for text-rich and broader VQA tasks when paired with efficient usage and integration strategies. Our code is publicly available at https://github.com/Linvyl/DAM-QA.git.
Sources
- Scalable Vision Language Model Training via High Quality Data Curation
- A Survey on LLM-as-a-Judge
- Multi-Agent VQA: Exploring Multi-Agent Foundation Models in Zero-Shot Visual Question Answering
- From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judge
- LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
- Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs
- Qwen2.5-VL Technical Report
- Tomayto, Tomahto. Beyond Token-level Answer Equivalence for Question Answering Evaluation
- Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
- Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
- Ovis: Structural Embedding Alignment for Multimodal Large Language Model
- ChartQAPro: A More Diverse and Challenging Benchmark for Chart Question Answering
- Enhancing Vietnamese VQA through Curriculum Learning on Raw and Augmented Text Representations
- Kosmos-2: Grounding Multimodal Large Language Models to the World
- VisualBERT: A Simple and Performant Baseline for Vision and Language
- UNIMO: Towards Unified-Modal Understanding and Generation via Cross-Modal Contrastive Learning
- Describe Anything: Detailed Localized Image and Video Captioning
- Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos
- TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models