ActiveSAM: Fast and Accurate Open-Vocabulary Semantic Segmentation with Frozen SAM 3
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "ActiveSAM: Fast and Accurate Open-Vocabulary Semantic Segmentation with Frozen SAM 3".
Jane: ActiveSAM introduces a training-free, zero-shot inference framework that adapts the frozen SAM 3 model into an active-vocabulary segmenter for open-vocabulary semantic segmentation (OVSS).
Tom: First, who's behind it and why it matters.
Title and authors: Tom: Moving on to the title and authors of "ActiveSAM: Fast and Accurate Open-Vocabulary Semantic Segmentation with Frozen SAM three" the name itself really highlights exactly what this paper is about—it’s about achieving both speed and accuracy in open-vocabulary segmentation while keeping SAM three frozen.
Jane: And it’s interesting to see that it's a training-free framework; that immediately tells us they aren't asking for massive labeled datasets to get this functionality working, which is always a huge win in the machine learning community.
Lu: The authors are clearly deep into the SAM family of models, which explains why they can leverage its frozen structure so effectively without needing to retrain the whole thing for every new segmentation task. They're building on something already very strong.
Meng: From an engineering angle, keeping a backbone frozen means we don't have to manage complex fine-tuning schedules or large memory footprints associated with updating millions of parameters, which simplifies the deployment pipeline significantly.
Lalam: It speaks to a maturing field where researchers are focusing on inference efficiency and practical utility rather than just pushing parameter counts higher in every model they build.
The paper's summary: Tom: So, the summary of ActiveSAM boils down to this: they introduce a zero-shot inference framework that transforms SAM three into an active-vocabulary segmenter by using class pruning guided by a low-resolution presence preview and then applying full-resolution decoding only to the relevant classes.
Jane: That’s a great way to put it simply; instead of running the heavy decoder on every single possible class name in the vocabulary, they use a quick look at the image first to decide which ones need detailed attention.
Lu: The technical mechanism involves several key components: they start with Contextual Prompt Expansion to enrich those raw class names, then use a preview stage at six hundred seventy-two by six hundred seventy-two resolution to estimate an image-conditioned active set, and finally use bucketed prompt multiplexing for the full-resolution decoding step.
Meng: I see the pipeline clearly now—it’s a multi-stage process where we filter down the work before committing to the most computationally intensive part of the model, which is exactly what we look for in optimizing inference.
Lalam: The summary really emphasizes that they are changing the computational structure so that it scales with how many classes are actually present in an image, not with the entire massive vocabulary size.
The paper's improvements: Tom: What really stands out about the improvements section is their focus on the technical mechanisms they built to make this work—specifically, Contextual Prompt Expansion for prompt enrichment and Margin-aware Background Calibration for final pixel refinement.
Jane: That makes sense; simply pruning classes isn't enough if the prompts themselves are weak or if you still get noisy background predictions at the final stage, which is where ActiveSAM goes beyond just a simple presence score filter.
Lu: The Contextual Prompt Expansion module is quite sophisticated because it doesn't just use the raw class name; it incorporates lexical canonicalization and retrieves semantic context from nearest neighbors and WordNet hypernyms to make the prompts much richer.
Meng: That caching of contextual information, retrieving those neighbors only once per dataset, is a smart way to ensure every prompt has that extra layer of semantic meaning without incurring massive per-image processing costs.
Lalam: It shows a deep understanding that high accuracy in OVSS isn't just about running the right classes; it’s about making sure the prompts for those classes are as informative as possible, which is really important for robust AI systems.
Conclusion: Tom: So, to wrap up on "ActiveSAM: Fast and Accurate Open-Vocabulary Semantic Segmentation with Frozen SAM three" the main conclusion is that this framework successfully adapts SAM three into an image-adaptive segmenter by leveraging early presence signals and active set decoding, leading to a performance gain of about +one point four mIoU over SegEarth-OV3 across eight benchmarks.
Jane: That performance jump is quite significant when you consider the speed they achieved, running up to five point five times faster on large datasets compared to previous state-of-the-art methods. It’s a solid trade-off for deployment in noisy environments.
Lu: What’s important here is that they managed to keep SAM three frozen and required no target dataset training, which keeps the model universally applicable across different tasks without needing specific fine-tuning for every new segmentation problem.
Meng: From a practical standpoint, this means we can deploy this kind of robust segmentation in scenarios where we need fast results without the overhead of constantly retraining or managing massive model updates.
Lalam: I think the implication for AI culture is that researchers are increasingly prioritizing efficiency and robustness alongside raw accuracy, which pushes the entire field toward building systems that are inherently practical for real-world use cases.
Tran Dinh Tien Zhiqiang Shen
Mohamed bin Zayed University of Artificial Intelligence
cs.CV, cs.AI, cs.LG
Submitted: 2026-06-15
Updated: 2026-09-29
Comments: Preprint. Code is available at https://github.com/VILA-Lab/ActiveSAM
Code: https://github.com/VILA-Lab/ActiveSAM
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 90/100
The gist: ActiveSAM introduces a training-free, zero-shot inference framework that adapts the frozen SAM 3 model into an active-vocabulary segmenter for open-vocabulary semantic segmentation (OVSS).
Key concepts
- ActiveSAM
- A training-free, zero-shot framework that adapts the frozen SAM 3 model into an active-vocabulary segmenter. It uses class pruning based on a low-resolution presence preview to apply full-resolution decoding only to relevant classes.
- Frozen SAM 3
- Leveraging the existing structure of the Segment Anything Model version 3 without needing to retrain the entire model. This keeps engineering simple by avoiding complex fine-tuning schedules and large memory footprints associated with updating millions of parameters.
- Contextual Prompt Expansion
- A module that enriches raw class names by incorporating lexical canonicalization and retrieving semantic context from nearest neighbors and WordNet hypernyms. This makes the prompts used for segmentation much richer.
- Bucket-based Prompt Multiplexing
- A technique used in the final decoding step where full-resolution decoding is applied only to the classes selected by the preview stage. This changes computational structure so it scales with present classes, not the entire vocabulary size.
Terminology
Summary
ActiveSAM introduces a training-free, zero-shot inference framework that adapts the frozen SAM 3 model into an active-vocabulary segmenter for open-vocabulary semantic segmentation (OVSS). This method addresses the inefficiency of running full-resolution decoding over an entire vocabulary when only a small subset of classes is relevant to a specific image. By leveraging SAM 3’s presence signal before full decoding and employing class pruning, ActiveSAM significantly improves the speed-accuracy tradeoff compared to existing state-of-the-art methods, achieving up to +1.4 mIoU improvement over SegEarth-OV3 while running up to 5.5× faster on large datasets, making it highly suitable for deployment in noisy, real-world domains.
Key Contributions and Framework Overview
ActiveSAM is a training-free inference framework that keeps all SAM 3 weights frozen, transforming the model into an image-adaptive segmenter at test time. The core idea is to use SAM 3’s presence signal before full-resolution decoding rather than only as a post-decoding filtering score. This changes the computational structure by invoking the expensive prompt-conditioned decoder only for an image-specific active vocabulary, which scales with the active set size rather than the full vocabulary size.
The framework consists of three main stages:
-
A low-resolution presence preview that estimates an image-conditioned active set, denoted as A(I).
-
Full-resolution decoding, where only prompts from the retained classes are decoded using bucketed prompt multiplexing.
-
Margin-aware background calibration to suppress low-confidence pixels and assign the background label when applicable.
Contextual Prompt Expansion (CPE)
To ensure robust alignment, ActiveSAM employs Contextual Prompt Expansion (CPE) to enrich raw class names with lexical and vocabulary context before decoding. This module performs two main tasks:
-
Lexical canonicalization: It applies a fixed, image-independent map, such as repairing compound labels like
windowpane7
towindow pane,
ensuring deterministic label artifacts are resolved before text encoding. -
Contextual enrichment: It retrieves class-level semantic context by finding the nearest other class names in the frozen pooled text-embedding space (Ms nearest neighbors) and incorporating WordNet noun hypernyms. The final prompt for each class is constructed by concatenating the canonicalized tokens with these cached neighbor and hypernym tokens, resulting in a contextual prompt, denoted as Z¯c.
Preview-driven Class Selection
The framework uses SAM 3’s low-resolution presence head as a preview to select the active vocabulary A(I). This stage is crucial for reducing latency by skipping unnecessary segmentation-head computation. The selection process involves:
-
A gating threshold, Vgate = 40, which bypasses the preview stage for small vocabularies (V ≤ Vgate).
-
For large vocabularies (V > Vgate), SAM 3 is run at a low resolution of rp = 672 to estimate a class-presence score qc for every class.
-
The final active set A(I) is determined by an image-adaptive quantile, defined as A(I) = (A0(I), A0(I) ≤ Vgate, arg TopKVgate c∈A0(I) qc).
Bucketed Full-Resolution Decoding and Calibration
Once the active set A(I) is selected, the full-resolution SAM 3 image encoder is evaluated once. The active classes are then partitioned into buckets (B1,..., BJ), where each bucket contains at most K=32 classes. The frozen grounding decoder processes these buckets using shared image features and the corresponding contextual prompts Z¯c. Scores from the instance and semantic heads are fused pixel-wise to obtain Sc(p) = max S inst c(p), Ssem c(p) for c ∈ A(I). Finally, Margin-aware Background Calibration (MABC) is applied:
-
It computes a margin score m(p) = s1(p) − s2(p), where s1 and s2 are the top two active-class scores at pixel p.
-
The MABC-calibrated confidence is cMABC(p) = s1(p) + max(s1 - s2, 0).
-
The final prediction yˆ(p) is set to background (bg) if cMABC(p) < tγ, where tγ is the dataset's background threshold and γ = 1.25.
Performance and Robustness
ActiveSAM achieves state-of-the-art performance across eight OVSS benchmarks, outperforming SegEarth-OV3 by +1.4 mIoU on average while running up to 5.5× faster on large datasets.
Improvements for AI systems
As a fastidious researcher, I have analyzed the ActiveSAM paper and identified several concrete, high-impact areas for improving current AI systems, particularly in the domain of Open-Vocabulary Semantic Segmentation (OVSS).
Here are the specific improvements and what the resulting system can achieve:
) 1. Inference Efficiency and Scalability Improvement
The most immediate improvement is a fundamental shift in computational architecture from full-vocabulary decoding to active-set decoding.
[Source: Abstract, Section 3.4, Computational Analysis]
ActiveSAM drastically reduces inference time by skipping the expensive full-resolution grounding decoder for classes not present in the image (via the low-resolution presence preview).
[Specific Improvement]: Implement a
presence-firstrouting mechanism where a frozen backbone's presence head dictates which prompts receive full computational load. This transforms prompt decoding from an operation scaling with vocabulary size to one scaling only with active class set size, achieving up to 5.5× faster inference on large datasets compared to SegEarth-OV3.
[System Capability]: Enables real-time or near real-time semantic segmentation for high-resolution imagery (e.g., autonomous driving feeds, high-throughput surveillance) where latency is critical, without sacrificing the accuracy achieved by the frozen SAM 3 backbone.
) 2. Robustness to Distribution Shift and Environmental Noise
The system explicitly shows superior performance when inputs are corrupted (noise, blur, compression), which is vital for real-world deployment.
[Source: Abstract, Section 4 (Robustness)]
ActiveSAM maintains high performance under Gaussian noise, motion blur, JPEG compression, and fog where CLIP-based methods degrade severely.
[Specific Improvement]: Integrate the margin-aware background calibration (MABC) mechanism as a primary noise filter during the final prediction stage. By using the margin between the top two scores to decide on background assignment rather than a simple absolute threshold, it enhances
local consistencyand fine-grained separation in complex scenes corrupted by visual artifacts.
[System Capability]: Creates more reliable perception systems for autonomous vehicles and embodied AI operating in unpredictable real-world conditions (e.g., adverse weather, sensor noise) by maintaining high mIoU stability compared to state-of-the-art baselines like CorrCLIP.
) 3. Enhanced Semantic Prompt Quality via Contextual Expansion
The method moves beyond raw class names by enriching them with lexical and hierarchical context before decoding.
[Source: Section 3.1, Contextual Prompt Expansion (CPE)]
ActiveSAM uses the frozen SAM 3 text encoder to generate prompts that are augmented by semantic-neighbor tokens
(nearest neighbors in embedding space) and WordNet hypernyms.
[Specific Improvement]: Systematically cache and retrieve class context (neighbor embeddings and WordNet hypernyms) once per dataset, ensuring every image prompt is rich with both immediate visual semantic context and broader hierarchical knowledge. This prevents poor alignment caused by isolated or ambiguous class names (e.g., correctly expanding
windowpaneinto related concepts likewindow pane).
[System Capability]: Improves the precision of segmentation in complex, fine-grained scenes (e.g., distinguishing between different types of furniture or architectural elements) where standard text prompts fail due to low semantic density.
) 4. Optimized Prompt Multiplexing and Grounding
The system uses bucketed prompt multiplexing to efficiently feed multiple active class prompts into the decoder in a single pass, rather than sequentially decoding each one.
[Source: Section 3.3, Bucketed Full-Resolution Decoding]
ActiveSAM groups active classes into buckets and feeds the shared image features and contextual prompts for that bucket simultaneously to the grounded decoder.
[Specific Improvement]: Implement dynamic prompt batching based on image content complexity or object density, rather than a fixed bucket size (K=32). This ensures that highly complex regions receive sufficient computational resources (shared cross-attention) without wasting cycles on simple regions.
[System Capability]: Increases throughput for high-density scenes by maximizing the efficiency of the frozen SAM 3 decoder, ensuring that the computational savings from pruning are fully realized in terms of speed and accuracy trade-off.
This set of improvements transforms a general OVSS pipeline into a highly efficient, robust, and contextually aware perception system capable of operating reliably in dynamic, real-world environments.
Abstract
Segment Anything Model 3 (SAM 3) provides a strong frozen backbone for concept-prompted segmentation, but applying it directly to open-vocabulary semantic segmentation (OVSS) is inefficient: full-resolution decoding is typically run over the entire dataset vocabulary, whereas each image contains only a small active subset of classes. We introduce ActiveSAM, a training-free inference framework that turns SAM 3 into an active-vocabulary segmenter. ActiveSAM first canonicalizes and expands class prompts, then uses evidence-proportional grounding to estimate an image-conditioned active set from a low-resolution presence preview. Only retained prompts receive full-resolution mask prediction, using bucketed prompt multiplexing with the frozen SAM 3 decoder. The preview stage uses only class-presence evidence and skips unnecessary segmentation-head computation. To resolve overlapping concept responses, exclusive concept decoding compares each pixel's joint score vector with class signatures estimated once per vocabulary from unlabeled images. ActiveSAM requires no weight updates, no oracle class-presence labels and no per-dataset hyperparameter tuning. Across eight OVSS benchmarks, ActiveSAM improves the speed-accuracy tradeoff of training-free open-vocabulary semantic segmentation, outperforming the current state-of-the-art SegEarth-OV3 by +2.1 mIoU on average while running much faster, with 7.3-12.2x speedups on large-vocabulary datasets. ActiveSAM also achieves the highest accuracy under image corruptions that simulate real-world distribution shift, making it well-suited for deployment in noisy-input domains such as autonomous driving and embodied AI. Code is available at https://github.com/VILA-Lab/ActiveSAM
Sources
- Perception Encoder: The best visual embeddings are not at the output of the network
- SAM 3: Segment Anything with Concepts
- Benchmarking Neural Network Robustness to Common Corruptions and Perturbations
- Language-driven Semantic Segmentation
- SegEarth-OV3: Exploring SAM 3 for Open-Vocabulary Semantic Segmentation in Remote Sensing Images
- OV-Stitcher: A Global Context-Aware Framework for Training-Free Open-Vocabulary Semantic Segmentation
- DINOv2: Learning Robust Visual Features without Supervision
- SAM 2: Segment Anything in Images and Videos
- MM-OVSeg:Multimodal Optical-SAR Fusion for Open-Vocabulary Segmentation in Remote Sensing
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models