SGP-SAM: Self-Gated Prompting for Transferring 3D Segment Anything Models to Lesion Segmentation

arXiv:2604.22825 · cs.CV, cs.AI · Submitted 2026-04-19 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "SGP-SAM: Self-Gated Prompting for Transferring 3D Segment Anything Models to Lesion Segmentation".

Tom: Large segmentation foundation models, such as Segment Anything Model (SAM), have advanced promptable segmentation in natural images,

Jane: First, who's behind it and why it matters.

Paper summary: Tom: So we've been deep in the weeds with SGP-SAM, and now we're wrapping up this segment by talking about what this whole paper means for segmentation in medical imaging.

Jane: Right, so to recap simply, SGP-SAM is a new way to take powerful three dee models that work on normal images and adapt them effectively for finding lesions by adding smart prompting and better loss functions.

Lu: It's fascinating how they're using the gating mechanism to decide when the model needs more spatial detail, which opens up so many creative avenues for how we structure these generative models.

Meng: I see the engineering challenge here is making sure that dynamic feature fusion actually runs smoothly and efficiently on real-time clinical hardware, which is something I'm focused on.

Lalam: From my perspective as an AI, this shows a cultural shift where we move toward models that aren't just pattern recognizers but are capable of contextually deciding when to enhance their understanding of complex visual data, which is really important for how we build trust in medical AI systems.

Tom: Exactly, and the authors, they’ve done a lot by showing consistent gains on datasets like the MSD Liver Tumor and Brain Tumor, proving that this transfer method outperforms standard fine-tuning approaches.

Jane: It really boils down to making segmentation more precise when the targets are tiny and buried in complex volumes, which is a huge practical win for diagnosis.

Lu: The way they combined Zoom Loss with the spatial gating shows a very elegant solution to those specific problems of small object supervision and extreme class imbalance.

Meng: For practical application, this means we can build tools that are significantly more reliable when identifying subtle abnormalities in scans, moving us closer to truly assistive diagnostics.

Lalam: I think the biggest impact is on how researchers approach model adaptation; it suggests we should focus less on brute-force fine-tuning and more on designing these intelligent prompting structures for domain-specific challenges.

Tom: So, to sum up, SGP-SAM is a framework that intelligently adapts foundation models for three dee lesion segmentation by using self-gated attention and targeted loss functions to boost spatial representation and focus on small areas.

Jane: It’s an important step in making AI tools robust enough for the nuanced reality of medical image analysis.

Lu: And it really shows how sophisticated prompting can be when you apply it to three dee data structures like this.

Meng: Next up, we'll take a look at how these specific results translate into real-world deployment scenarios for clinical support systems.

Conclusion: Tom: So, we've been walking through the technical details of SGP-SAM, and now we're focusing on what this whole paper really means for its title and who came up with it.

Jane: Right, so basically, they’re introducing a new method called SGP-SAM that uses self-gated prompting to adapt models like SAM for three dee lesion segmentation.

Lu: It's interesting how the title itself highlights the core mechanism: self-gated prompting—the model making its own decisions about needing spatial enhancement dynamically. That’s a very creative approach to prompt engineering in a medical context.

Meng: From an engineering viewpoint, focusing on the authors and their specific choices in architecture is important because it tells us exactly where the innovation lies for implementation.

Lalam: I think it shows a major cultural shift where we move toward models that aren't just pattern recognizers but are capable of contextually deciding when to enhance their understanding of complex visual data, which is really important for how we build trust in medical AI systems.

Tom: Exactly, and the authors clearly showed they understood the specific spatial challenges of three dee volumes, moving beyond standard image segmentation techniques.

Jane: It boils down to making sure these powerful foundation models can handle the unique demands of three dee medical data without needing a complete rewrite every time.

Lu: The way they designed Zoom Loss specifically for lesion-focused supervision is a testament to their deep understanding of class imbalance in this domain. That’s where the real theoretical beauty is.

Meng: I'm thinking about the practical impact on development pipelines; if these methods work consistently, it means we can move toward faster iteration cycles for clinical tools.

Lalam: I see this as an advance in how researchers approach model adaptation; it suggests we should focus less on brute-force fine-tuning and more on designing these intelligent prompting structures for domain-specific challenges.

Tom: So, to wrap up this discussion on SGP-SAM: it’s a framework that uses self-gated prompting and targeted loss functions to enhance three dee SAM models for lesion segmentation by improving spatial representation and focusing on small areas. Jane, Lu, Meng, Lalam—thanks for joining us on this deep dive into the paper.

Jane: It’s an important step in making AI tools robust enough for the nuanced reality of medical image analysis.

Lu: And it really shows how sophisticated prompting can be when you apply it to three dee data structures like this.

Meng: Next up, we'll take a look at how these specific results translate into real-world deployment scenarios for clinical support systems.

School of Intelligent Systems Engineering, Sun Yat-sen University

cs.CV, cs.AI

Submitted: 2026-04-19

Updated: 2026-09-28

Importance score: 76/100

The gist: Large segmentation foundation models, such as Segment Anything Model (SAM), have advanced promptable segmentation in natural images, but directly transferring these 3D SAM-style models to lesion

Key concepts

Self-Gated Prompting Module (SGPM)
This module sits within the 3D image encoder and learns dynamically when intermediate feature maps require spatial enhancement. It uses a gating unit to decide if fusion is needed, activating a multi-scale block only when necessary to improve feature representation.
Multi-Scale Feature Fusion Block (MSFB)
When activated by the SGPM, this block extracts information from different scales of features. It compresses the channels first and then uses three convolutional branches with different kernel sizes (1x1x1, 3x3x3, 5x5x5) to capture multi-scale spatial details before fusing them back together.
Zoom Loss
This loss function is designed to handle small lesion areas and class imbalance. It combines Dice loss with a voxel-balanced focal term and a size-dependent reweighting factor, forcing the model to pay more attention to smaller lesions instead of easy background pixels.

Terminology

Summary

Large segmentation foundation models, such as Segment Anything Model (SAM), have advanced promptable segmentation in natural images, but directly transferring these 3D SAM-style models to lesion segmentation remains challenging due to weak spatial representational capacity for small targets and extreme foreground–background imbalance in 3D volumes.

The gist

SGP-SAM is a self-gated prompting framework designed for efficient and effective transfer to 3D lesion segmentation by performing conditional multi-scale spatial enhancement when features require it, while using Zoom Loss to strengthen lesion-focused supervision.

How it works

The core of SGP-SAM is the Self-Gated Prompting Module (SGPM), which is inserted into the 3D image encoder to learn when a feature map needs additional multi-scale fusion. This module consists of two main components:

  1. A Multi-Channel Self-Gating Unit that estimates whether intermediate features need additional spatial enhancement. This unit computes dimension-wise summary vectors by mean pooling along complementary axes and maps these summaries to a scalar key, which is then aggregated using learnable weights normalized by softmax to produce a gate logit, denoted as 's'.

  2. A Multi-Scale Feature Fusion Block (MSFB) that performs multi-scale fusion when activated. This block compresses the channels first using a 1x1x1 convolution and then applies three convolutional branches of different kernel sizes (1x1x1, 3x3x3, and 5x5x5) to extract multi-scale information. These outputs are fused by averaging before a final re-expansion through a 1×1×1 convolution.

The output feature map is then determined by the gate value 'g' (derived from the Gumbel-Sigmoid estimator) via the formula:

Fout = g · MSFB(F) + (1 − g) · F. If features are predicted to be spatially insufficient, the MSFB enhances them; otherwise, features pass through unchanged.

Zoom Loss for Small Lesions

To address the challenge of small lesion areas in 3D medical images and extreme class imbalance, SGP-SAM introduces a Zoom Loss. This loss strengthens lesion-focused supervision by combining Dice loss with a voxel-balanced focal term and incorporating a lesion-size-dependent reweighting factor. The final loss formulation is:

LV = − (1/N) Σ i=1 [αyi(1 − pi)γ log(pi) + (1 − α)(1 − yi)pγi log(1 − pi)]. This mechanism encourages the model to focus on smaller and more complex lesion areas, mitigating the influence of easy background pixels.

Key Contributions and Results

The paper's primary contributions are:

– Proposing SGP-SAM, a self-gated prompting framework for transferring 3D SAM-style models to lesion segmentation with improved spatial representation and lesion-focused learning.

– Designing the SGPM, where a multi-channel gating unit predicts when to activate an MSFB for conditional spatial enhancement inside the 3D image encoder.

– Introducing Zoom Loss, which combines Dice Loss and voxel-balanced focal supervision with lesion-size-aware weighting to handle small lesion areas effectively.

Experiments on MSD Liver Tumor and MSD Brain Tumor (enhancing tumor) show consistent gains over strong transfer baselines based on SAM-Med3D. For the MSD Liver Tumor, SGP-SAM improves mDice by 7.3% over fine-tuning, reaching a mDice of 0.7151 compared to 0.6667 for fine-tuning alone. The ablation study confirms that combining SGPM and Zoom Loss further improved performance, with the final results achieving an mIoU of 0.5775 and an mDice of 0.7151 on the MSD Liver Tumor dataset. The optimal placement for SGPM was found to be at the end of each block in the image encoder.

Comparison with Related Methods

The comparative experiments demonstrate that SGP-SAM outperforms baselines like pre-trained SAM-Med3D enhanced through fine-tuning. For the MSD Liver Tumor dataset, SGP-SAM achieved an increase of 7.8% in mIoU and 7.3% in mDice compared to the fine-tuning method, and a 4.9% improvement in mIoU and 4.2% improvement in mDice compared to fine-tuning on the MSD Brain Tumor dataset. Visualization results confirm that for liver tumors with small proportions, Zoom Loss allows the model to pay more attention to smaller areas, while SGPM plays a pivotal role for brain tumors with complex structures by enriching the representation of features at multiple scales.

Ablation Studies

Ablation results on the MSD Liver Tumor dataset show the effectiveness of both SGPM and Zoom Loss. Adding SGPM increased mIoU to 0.

Improvements for AI systems

Here are the specific improvements that can be made to existing medical image segmentation AI systems by implementing the SGP-SAM framework, and what those improved systems will be capable of doing:

  1. Enhance Spatial Representation in Small/Irregular Targets: The system will gain a more nuanced understanding of small, irregular lesion boundaries (e.g., early-stage liver tumors or subtle enhancing tumor margins) because the Self-Gated Prompting Module (SGPM) dynamically activates multi-scale feature fusion only when features are deemed spatially insufficient.

  2. Improve Efficiency by Avoiding Unnecessary Computation: The system will operate more efficiently than models that always perform expensive multi-scale fusion, as the SGPM selectively applies this enhancement only to challenging regions, saving computational resources on easier background or large structures.

  3. Boost Performance on Extreme Class Imbalance: By incorporating the Zoom Loss (combining Dice Loss and a voxel-balanced focal term with lesion-size-dependent reweighting), the system will significantly reduce overfitting to easy background voxels. This allows it to achieve higher accuracy (up to 7.3% gain in mDice on Liver Tumor) specifically for small, critical lesion areas.

  4. Enable Robust Transfer Learning from General Models: The SGP-SAM framework provides a structured and effective way to transfer the powerful promptable capabilities of large foundation models like SAM (and its medical extensions, e.g., SAM-Med3D) to specialized 3D lesion segmentation tasks, overcoming the limitations of direct transfer.

  5. Handle Complex 3D Anatomical Structures: The combination of multi-scale feature fusion via MSFB and the self-gated mechanism allows the system to effectively segment lesions with complex shapes and ambiguous boundaries (as demonstrated on MSD Brain Tumors).

  6. Achieve State-of-the-Art Segmentation Metrics: The resulting AI systems are expected to achieve superior segmentation metrics (mIoU and mDice) compared to standard fine-tuning baselines, specifically showing gains of up to 7.3% in Dice score on liver tumor segmentation tasks.

Sources

Related papers