From Reasoning to Pixels: Grounded Medical Multimodal LLMs for VQA and Segmentation
cs.CV, cs.AI
Submitted: 2026-08-27
Updated: 2026-08-27
Comments: accepted by ECCV 2026
Code: https://github.com/NUST-Machine-Intelligence-Laboratory/MedREAL
License: http://creativecommons.org/licenses/by/4.0/
The gist: Although Multimodal Large Language Models (MLLMs) have demonstrated impressive performance in Medical Visual Question Answering (Med-VQA), their reliance on global image features often lacks precise
Terminology
Abstract
Although Multimodal Large Language Models (MLLMs) have demonstrated impressive performance in Medical Visual Question Answering (Med-VQA), their reliance on global image features often lacks precise pixel-level grounding, thereby limiting clinical trustworthiness. To bridge the semantic gap between high-level clinical reasoning and spatial localization, we propose MedREAL (Medical REasoning-driven Answering and Localization), a unified framework that seamlessly aligns linguistic reasoning with spatial grounding. Specifically, MedREAL introduces Seg Anchored Reasoning Pooling (SARP) to distill task-relevant semantic evidence directly from[SEG] tokens within the MLLM's hidden states. Furthermore, a Reasoning-to-Visual (R2V) fusion mechanism is proposed to effectively inject these reasoning-aware features into a segmentation pipeline for accurate mask decoding. To facilitate this paradigm, we construct MedRAVS-13K, a comprehensive dataset comprising 13,824 expertly validated samples across four diverse imaging modalities. Extensive experiments demonstrate that MedREAL significantly outperforms state-of-the-arts, achieving 68.49% gIoU and 70.47% cIoU on benchmark evaluations. By generating evidence masks that are strictly consistent with textual diagnoses, MedREAL provides a robust, interpretable framework for reasoning-driven medical image analysis.
Sources
- GPT-4 Technical Report
- Qwen Technical Report
- Qwen3-VL Technical Report
- Skin Lesion Analysis Toward Melanoma Detection 2018: A Challenge Hosted by the International Skin Imaging Collaboration (ISIC)
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- PathVQA: 30000+ Questions for Medical Visual Question Answering
- SEG-SAM: Semantic-Guided SAM for Unified Medical Image Segmentation
- Medical SAM3: A Foundation Model for Universal Prompt-Driven Medical Image Segmentation
- Segment Anything
- MedSAM3: Delving into Segment Anything with Medical Concepts
- Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement
- Decoupled Weight Decay Regularization
- ISLES'24 -- A Real-World Longitudinal Multimodal Stroke Dataset
- DeepSeek-OCR: Contexts Optical Compression
- Medical SAM 2: Segment medical images as video via Segment Anything Model 2
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models