MedSAM3: Delving into Segment Anything with Medical Concepts
cs.CV, cs.AI
Submitted: 2025-11-24
Updated: 2026-09-14
Code: https://github.com/Joey-S-Liu/MedSAM3
License: http://creativecommons.org/licenses/by/4.0/
The gist: Medical image segmentation is fundamental for biomedical discovery.
Terminology
Abstract
Medical image segmentation is fundamental for biomedical discovery. Existing methods lack generalizability and demand extensive, time-consuming manual annotation for new clinical application. Here, we propose MedSAM-3, a text promptable medical segmentation model for medical image and video segmentation. By fine-tuning the Segment Anything Model (SAM) 3 architecture on medical images paired with semantic conceptual labels, our MedSAM-3 enables medical Promptable Concept Segmentation (PCS), allowing precise targeting of anatomical structures via open-vocabulary text descriptions rather than solely geometric prompts. We further introduce the MedSAM-3 Agent, a framework that integrates Multimodal Large Language Models (MLLMs) to perform complex reasoning and iterative refinement in an agent-in-the-loop workflow. Comprehensive experiments across diverse medical imaging modalities, including X-ray, MRI, Ultrasound, CT, and video, demonstrate that our approach significantly outperforms existing specialist and foundation models. We will release our code and model at https://github.com/Joey-S-Liu/MedSAM3.
Sources
- M3D: Advancing 3D Medical Image Analysis with Multi-Modal Large Language Models
- TransUNet: Transformers Make Strong Encoders for Medical Image Segmentation
- Skin Lesion Analysis Toward Melanoma Detection 2018: A Challenge Hosted by the International Skin Imaging Collaboration (ISIC)
- ISLES'24: Final Infarct Prediction with Multimodal Imaging and Clinical Data. Where Do We Stand?
- Polyp-PVT: Polyp Segmentation with Pyramid Vision Transformers
- MedRAX: Medical Reasoning Agent for Chest X-ray
- MDocAgent: A Multi-Modal Multi-Agent Framework for Document Understanding
- MMedAgent: Learning to Use Medical Tools with Multi-modal Agent
- Language-guided Medical Image Segmentation with Target-informed Multi-level Contrastive Alignments
- Efficient automatic segmentation for multi-level pulmonary arteries: The PARSE challenge
- U-Mamba: Enhancing Long-range Dependency for Biomedical Image Segmentation
- MedSAM2: Segment Anything in 3D Medical Images and Videos
- Attention U-Net: Learning Where to Look for the Pancreas
- SAM 2: Segment Anything in Images and Videos
- InternAgent: When Agent Becomes the Scientist -- Building Closed-Loop System from Hypothesis to Verification
- The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)
- A Survey on Agentic Multimodal Large Language Models
- ChemLLM: A Chemical Large Language Model
- BiomedParse: a biomedical foundation model for image parsing of everything everywhere all at once
- Medical SAM 2: Segment medical images as video via Segment Anything Model 2
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models