MedVL-SAM2: A unified 3D medical vision-language model for multimodal reasoning and prompt-driven segmentation

arXiv:2601.09879 · cs.CV, cs.AI · Submitted 2026-01-14 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "MedVL-SAM2: A unified 3D medical vision-language model for multimodal reasoning and prompt-driven segmentation".

Jane: MedVL-SAM2 introduces a unified 3D medical multimodal model that concurrently supports report generation, VQA, and multiparadigm segmentation, including semantic, referring, and interactive segmentation.

Tom: First, who's behind it and why it matters.

Title and authors: Tom: So, let's get into the specifics of this paper now that we have you all tuned in. This paper is titled "MedVL-SAM2: A unified three dee medical vision–language model for multimodal reasoning and prompt-driven segmentation," and it was authored by Yang Xing, Jiong Wu, Savas Ozdemir, Ying Zhang, Yang Yang, Wei Shao, and Kuang Gong. It’s a big team coming together to tackle this specific problem of integrating language understanding with three dee medical data analysis.

Jane: That title really sums up what they are trying to achieve: unifying multimodal reasoning across different tasks like segmentation and question answering in a three dee medical context using prompt-driven methods. It suggests they aren't just building another image analyzer, but something that understands the underlying structure of the three dee data while also talking about it.

Lu: The authors are from some very strong institutions, which I think is telling because they have the resources to tackle a problem this deep within medical imaging. Their focus on unifying reasoning and segmentation across multiple modalities is exactly where cutting-edge AI research needs to go in healthcare right now.

Meng: I see the ambition there, but I wonder how much of that unification actually translates into something practically deployable for a hospital setting, especially when you're dealing with complex three dee scans where errors can be quite costly.

Lalam: What’s exciting is that they are tackling the problem head-on by proposing this unified model to handle report generation, VQA, and multiple segmentation paradigms simultaneously. That breadth of capability is what makes me think this work could have a huge impact on how clinicians interact with their scans.

The paper's summary: Tom: So, summarizing the core idea of MedVL-SAM2, the paper proposes a single three dee medical multimodal model designed to perform report generation, visual question answering, and segmentation in various forms—semantic, referring, and interactive. Essentially, they are trying to bridge the gap between understanding complex medical images from a language prompt and performing detailed pixel-level tasks on those images.

Jane: To put it simply for our listeners, imagine you can ask a complex question about a three dee CT scan, get a detailed text report back, and then tell the AI exactly which specific structure in that scan you want it to highlight or segment using simple instructions like pointing or drawing a box. That’s the main concept here.

Lu: They achieve this by integrating image-level reasoning with pixel-level perception through a cohesive architecture specifically built for three dee medical imaging, utilizing a SAM2-based volumetric segmentation module to get precise multi-granular spatial reasoning across the volume.

Meng: I read that they train it in a multi-stage pipeline: first pretraining on three dee CT image–text pairs to align visual features with radiology language embeddings, and then jointly optimizing it with language understanding and segmentation objectives using a comprehensive three dee CT dataset. That seems like a solid, if lengthy, training strategy.

Lalam: That joint training approach is key because it allows for flexible interaction through language prompts, point prompts, or bounding-box prompts to unify the high-level reasoning with accurate spatial localization. It makes the model much more versatile than previous models that were stuck in one mode or another.

The paper's improvements: Tom: Moving on to what they actually improved, the paper highlights several architectural choices made in MedVL-SAM2 to solve specific problems inherent in previous three dee medical VLMs. They introduced a "three dee-aware vision encoder" called Mthree dee-CLIP to better preserve volumetric context, and they also addressed computational efficiency by using an MLP Mixer projection layer that reduces tokens down to five hundred twelve while keeping the spatial information intact.

Jane: That token compression is interesting; reducing the input dimension from two thousand forty-eight tokens to five hundred twelve using that MLP-Mixer layer helps manage the computational load significantly without losing the crucial spatial details needed for medical analysis.

Lu: They also tackled cross-slice consistency by using the memory attention mechanism from SAM2 to ensure that there's continuity maintained across all slices within the volume, which is something 2D slice-based analyses always struggle with.

Meng: From an engineering viewpoint, managing that token reduction while preserving spatial information is a delicate balance. If you lose too much detail during compression, the downstream tasks like segmentation could become inaccurate in critical areas.

Lalam: The use of SAM2 as the segmentation module is a major improvement because it gives them strong performance and native support for multi-prompt interaction, which directly addresses that need for precise pixel-level grounding we talked about earlier.

Conclusion: Tom: So, to wrap up this discussion on MedVL-SAM2, the main takeaway is that this unified three dee medical vision–language model successfully combines report generation and VQA with various forms of segmentation—semantic, referring, and interactive—all while maintaining volumetric context through clever architectural choices.

Jane: It really seems like the framework’s strength lies in its ability to handle these diverse tasks coherently because of that unified approach, allowing it to leverage language prompts effectively for spatial grounding in three dee data.

Lu: The implications here are huge for how we can interpret complex medical scans, moving beyond just generating text reports to actually performing precise structural analysis on the volume itself.

Meng: I think the practical impact will be seen as systems that can help radiologists and clinicians perform more detailed, structure-specific assessments directly on the data in a more interactive way.

Lalam: This work really sets a high bar for future models because it shows how to integrate visual reasoning and spatial grounding into one cohesive system, which could improve the way we train and deploy AI for clinical support.

Department of Biomedical Engineering, University of Florida, Gainesville, FL, USA · Department of Radiology, University of Florida, Jacksonville, FL, USA · Research Computing, University of Florida, Gainesville, FL, USA · Department of Medicine, University of Florida-Gainesville

cs.CV, cs.AI

Submitted: 2026-01-14

Updated: 2026-10-01

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 90/100

The gist: MedVL-SAM2 introduces a unified 3D medical multimodal model that concurrently supports report generation, VQA, and multiparadigm segmentation, including semantic, referring, and interactive

Key concepts

Unified 3D Medical Multimodal Model
This is a single architecture designed to handle different medical tasks simultaneously. It connects high-level reasoning (understanding text and images) with low-level perception (identifying specific pixels in a 3D scan). This allows the model to perform complex jobs like writing reports, answering questions, and precisely segmenting organs within 3D scans.
SAM2-based Volumetric Segmentation Module
This module uses the SAM2 framework to provide precise segmentation capabilities across a full 3D volume. It leverages a special token called '[SEG]' to receive semantic prompts from the language model, enabling it to generate accurate masks for different structures, supporting semantic, referring, and interactive segmentation.
Promptable Segmentation
The model can be guided by three different types of inputs: natural language (text), point locations (specific 3D coordinates), and bounding boxes. This flexibility allows users to interact with the model in various ways to guide the pixel-level segmentation, leading to better results, especially when using bounding-box prompts.
3D-aware Vision Encoder (M3D-CLIP)
This component is responsible for understanding the 3D medical images. It uses a CLIP-based encoder adapted for 3D data to capture the volumetric context of CT scans effectively. This ensures that the model understands not just individual slices, but how different parts of the entire 3D volume relate to each other, preserving crucial spatial information.

Terminology

Summary

MedVL-SAM2 introduces a unified 3D medical multimodal model that concurrently supports report generation, VQA, and multiparadigm segmentation, including semantic, referring, and interactive segmentation. This model unifies image-level reasoning with pixel-level perception through a cohesive architecture tailored for 3D medical imaging.

How it works

The proposed MedVL-SAM2 integrates image-level reasoning and pixel-level perception through a cohesive architecture tailored for 3D medical imaging, incorporating a SAM2-based volumetric segmentation module to enable precise multi-granular spatial reasoning. The model is trained in a multi-stage pipeline: it is first pretrained on a large-scale corpus of 3D CT image–text pairs to align volumetric visual features with radiology-language embeddings. It is then jointly optimized with both language understanding and segmentation objectives using a comprehensive 3D CT segmentation dataset. This joint training enables flexible interaction via language, point, or box prompts, thereby unifying high-level visual reasoning with spatially precise localization.

Architecture Overview

The architecture consists of a LLaVA-like VLM and a promptable segmentation module. The VLM comprises: (i) a volumetric vision encoder, (ii) an adaptive token projection layer, and (iii) an LLM backbone with LoRA layers. To address three key challenges in unified 3D medical VLM design:

  1. Volumetric Understanding: Addressed through a 3D-aware vision encoder (M3D-CLIP) that preserves volumetric context.

  2. Computational Efficiency: Addressed via an MLP Mixer projection layer that reduces tokens to 512 while preserving spatial information.

  3. Cross-Slice Consistency: Alleviated by leveraging the memory attention mechanism of SAM2 to maintain slice-wise consistency throughout the volume.

Vision-Language Processing Pipeline

The volumetric vision encoder utilizes a CLIP-based 3D encoder, deriving image embeddings where vin = Evc(Iin) ∈ R n×d, with visual tokens derived from a pre-trained 3D Vision Transformer (3D ViT). The adaptive token projection layer compresses the high dimensionality of 3D medical images, reducing the token count from 2,048 to 512 using an MLP-Mixer layer formulated as equation (2). The LLM backbone is InternVL2.5-4B [6], utilizing LoRA layers for efficient fine-tuning.

[SEG] Token Controlled Segmentation

SAM2 is employed as the segmentation module due to its strong performance and native support for multi-prompt interaction. Pixel-level grounding is enabled through a hidden state of the [SEG] token, denoted as [SEG]hs, which serves as a semantic prompt. This hidden state is passed to the SAM2 decoder, where it is concatenated with prompt embeddings from the SAM2 prompt encoder and decoded into segmentation masks: Mout = SAM2(Iin, [SEG]hs,Pin). During training, gradients flowing through the [SEG] token enable the VLM to learn how to use it for accurate grounding.

Key Contributions and Evaluation

The key contributions include providing a unified 3D medical VLM capable of jointly performing image-level reasoning and pixel-level perception, supporting report generation, VQA, and multiple types of 3D segmentation. The framework incorporates multi-modal prompting through language, point, and bounding-box prompts. Performance comparisons show superior results across tasks:

- Report Generation:

The model achieves a score of 89.33 on the CT-RATE dataset for report generation, surpassing M3D and achieving comparable or higher scores than the current state-of-the-art CT-Chat.

- VQA:

The model achieves 89.74% accuracy on multiple-choice VQA, surpassing CT-Chat by a clear margin.

- Referring Segmentation:

The model achieves the highest Dice scores in both referring and semantic segmentation tasks compared to SegVol and M3D, demonstrating stronger spatial reasoning and text-conditioned localization abilities.

- Interactive Segmentation:

Bounding-box prompts yield the highest Dice scores, followed by point prompts. This confirms that incorporating user feedback substantially improves segmentation quality, especially for challenging structures like the kidneys and pancreas.

Training Strategy

The model employs a three-stage progressive training strategy inspired by Cambrian-1 [40] framework to stably align 3D visual features, language representations, and spatial grounding:

  1. Stage 1: Optimizing only the MLP-Mixer projection layer for vision-language alignment (3 epochs on report generation subset).

  2. Stage 2: Expanding trainable set to include the vision encoder, projection layer, and LLM LoRA parameters (5 epochs on CT-RATE for report generation and VQA).

Improvements for AI systems

As a fastidious and diligent researcher, I have analyzed the architecture, training methodology, and experimental results of MedVL-SAM2. To improve AI systems using this scientific paper—specifically in the domain of 3D medical imaging—the following targeted enhancements should be implemented:


The core improvements focus on enhancing the model's capabilities across its three unified tasks: Report Generation (RG), Visual Question Answering (VQA), and Multi-Granular Segmentation (Semantic, Referring, Interactive).

Here are specific improvements and the resulting enhanced AI system capabilities:

  1. Enhance the 3D Volumetric Understanding via Encoder Selection:

  2. Refine Token Compression for Efficiency:

  3. Integrate Cross-Modal Consistency via Advanced Memory Attention:

  4. Optimize Multi-Stage Training Stability and Convergence:

By implementing these improvements, the resulting AI system will achieve the following specific capabilities:

  1. The improved system will possess superior ability to generate highly accurate, clinically nuanced radiology reports (RG), capable of distinguishing subtle pathological findings from normal anatomy with high semantic fidelity (as demonstrated by outperforming CT-Chat in both BLEU and ROUGE scores).

  2. It will enable precise, pixel-level localization of complex 3D anatomical structures (e.g., liver, kidney, pancreas) through semantic segmentation, achieving state-of-the-art Dice scores across various benchmarks (as shown by outperforming M3D and SegVol).

  3. The system will exhibit robust ability to perform fine-grained spatial reasoning via referring segmentation, allowing it to accurately segment structures based on implicit textual descriptions rather than explicit class names (as evidenced by competitive performance against M3D in referring segmentation).

  4. It will gain the critical capability of interactive 3D segmentation, enabling a human-in-the-loop workflow where users can iteratively refine masks using point or bounding-box prompts, providing significant gains for complex structures like small vessels and ribs (as shown by superior Dice scores with interactive prompts on the TotalSegmentor dataset).

  5. The system will demonstrate enhanced cross-modal reasoning in VQA, allowing it to provide accurate answers to complex medical questions by synthesizing both high-level clinical narratives and precise 3D spatial grounding, leading to higher accuracy in multiple-choice and long-answer VQA tasks.

Abstract

Recent progress in medical vision-language models (VLMs) has achieved strong performance on image-level text-centric tasks such as report generation and visual question answering (VQA). However, achieving fine-grained visual grounding and volumetric spatial reasoning in 3D medical VLMs remains challenging, particularly when aiming to unify these capabilities within a single, generalizable framework. To address this challenge, we proposed MedVL-SAM2, a unified 3D medical multimodal model that concurrently supports report generation, VQA, and multi-paradigm segmentation, including semantic, referring, and interactive segmentation. MedVL-SAM2 integrates image-level reasoning and pixel-level perception through a cohesive architecture tailored for 3D medical imaging, and incorporates a SAM2-based volumetric segmentation module to enable precise multi-granular spatial reasoning. The model is trained in a multi-stage pipeline: it is first pre-trained on a large-scale corpus of 3D CT image-text pairs to align volumetric visual features with radiology-language embeddings. It is then jointly optimized with both language-understanding and segmentation objectives using a comprehensive 3D CT segmentation dataset. This joint training enables flexible interaction via language, point, or box prompts, thereby unifying high-level visual reasoning with spatially precise localization. Our unified architecture delivers state-of-the-art performance across report generation, VQA, and multiple 3D segmentation tasks. Extensive analyses further show that the model provides reliable 3D visual grounding, controllable interactive segmentation, and robust cross-modal reasoning, demonstrating that high-level semantic reasoning and precise 3D localization can be jointly achieved within a unified 3D medical VLM.

Sources

Related papers