AD-SAM: Adapting the Segment Anything Model for Semantic Segmentation in Autonomous Driving
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "AD-SAM: Adapting the Segment Anything Model for Semantic Segmentation in Autonomous Driving".
Jane: AD-SAM presents a fine-tuned vision foundation model designed for semantic segmentation in autonomous driving, significantly enhancing performance over existing models by integrating dual encoders and a hybrid loss function.
Tom: First, who's behind it and why it matters.
Title and authors: Tom: So we’ve got the title "AD-SAM: Adapting the Segment Anything Model for Semantic Segmentation in Autonomous Driving," and it really highlights that they are taking a general foundation model, SAM, and adapting it specifically for autonomous driving segmentation tasks. It tells us exactly what the focus of this work is all about.
Jane: Exactly; it points out that they aren't just using SAM out of the box but are actively fine-tuning it to address the unique spatial and geometric challenges that exist when a car has to perceive its surroundings on a road. It’s about making sure the segmentation works perfectly for driving, not just general images.
Lu: The authors clearly want to show how this adaptation improves performance in scenarios where standard segmentation methods might struggle with the specific visual characteristics of driving environments, which is a key focus area for this kind of work.
Meng: I’m wondering about the authors themselves; are they from a team that has direct experience with autonomous systems, or is it purely a vision-focused group applying their knowledge to the AD domain? That background can really influence how practical the results are.
Lalam: The fact that they built something specifically for autonomous driving suggests an understanding of the safety-critical nature of segmentation in this field, which I think is important because it means we're looking at applications where precision has direct real-world consequences.
The paper's summary: Tom: Moving on to the summary, the AD-SAM architecture essentially builds on SAM by adding a dual-encoder system—one for broad semantic context and one for fine local detail—and then uses a deformable decoder to progressively refine those features into final class predictions. It’s like giving the model two different ways to look at an image simultaneously.
Jane: That dual-encoder setup is pretty intuitive; having one part handle the big picture context, like knowing there's a road ahead, and another part focusing on the small details, like exactly where a curb starts or ends. The summary emphasizes that this fusion module is key to aligning those different feature types across various scales and object shapes.
Lu: The paper highlights that this multi-scale feature fusion, which involves computing offset fields and modulation masks using an equation like y(p2) = 5w1/three hundred fourteen/cdot x(p2 + p1 + Δp1) ⋅ m1, is what allows the model to adaptively transform the input features based on where they need to look most closely.
Meng: From an engineering standpoint, that deformable fusion module sounds like it adds complexity during inference, so I’m wondering how they manage that computational overhead when you're running this on a car's onboard computer versus just testing it on a powerful GPU.
Lalam: I think the summary really emphasizes the hybrid loss function they use—integrating Focal loss, Dice loss, LovászSoftmax loss, and Surface loss—to ensure the model gets good results in terms of class balance and sharp boundaries simultaneously. That shows a deep consideration for segmentation quality across different metrics.
The paper's improvements: Tom: The paper outlines several specific enhancements they made to SAM, primarily focusing on that dual-encoder structure and the multi-stage decoder that uses deformable attention to refine the features progressively through three distinct stages. That sequential refinement seems like a sophisticated way to build up the final segmentation mask.
Jane: Those three stages—Stage one at one thousand twenty-four channels down to Stage three at just sixty-four channels—suggest a very careful process of progressively reducing complexity while maintaining accuracy, which is smart when you need precise boundaries for things like road edges.
Lu: The hybrid loss function itself, combining Focal loss with Dice loss and Surface loss, is designed to tackle several issues at once: imbalance in classes, overlap optimization via IoU using Dice, boundary accuracy using surface losses based on distance transforms, and stability through LovászSoftmax.
Meng: I’m interested in the training specifics; they mention training on Cityscapes with Focal loss and BDD100K with Dice loss and Surface loss. How did they balance those different optimization targets during the one hundred epochs of training?
Lalam: The paper points out that this combination of losses is specifically chosen to improve semantic class balance, boundary precision, and optimization stability, which is crucial when dealing with the inherent complexities found in driving scenes where everything from a tiny pothole to a large vehicle needs careful delineation.
Conclusion: Tom: So, wrapping up the discussion on "AD-SAM: Adapting the Segment Anything Model for Semantic Segmentation in Autonomous Driving," it’s clear they’ve taken SAM and engineered it with dual encoders and a deformable decoder, guided by a hybrid loss function to get superior segmentation performance on both Cityscapes and BDD100K datasets.
Jane: The main implication here is that this model shows how foundation models can be effectively specialized for highly specific, complex domains like autonomous driving perception, demonstrating strong data efficiency because it learns well even with relatively little labeled supervision compared to older methods.
Lu: The results show AD-SAM achieving sixty-eight point one four mIoU on Cityscapes and fifty-nine point five zero mIoU on BDD100K, while also showing faster learning dynamics and better cross-domain retention scores when compared against models like SAM or G-SAM.
Meng: For practical implementation, the fact that it converges quickly in about thirty to forty epochs is a big plus for getting iterative improvements going in a real engineering pipeline; it suggests we might be able to get usable results without needing massive annotation efforts upfront.
Lalam: I think the overall impact of AD-SAM is that it pushes the envelope on how we can leverage pre-trained vision models to achieve high reliability in perception systems, which could eventually lead to safer and more intuitive AI agents interacting with our physical world.
Tom: Absolutely, so that's the AD-SAM story—a tailored approach using dual encoders and sophisticated loss balancing that gives us a much better segmentation tool for autonomous driving perception than what we had before. Next time we talk about papers, we’ll see how this work compares to other approaches tackling domain shift in vision systems.
Mario Camarena, Het Patel, Fatemeh Nazari, Evangelos Papalexakis, Mohamadhossein Noruzoliaee, Jia Chen
Department of Computer Science, The University of Texas Rio Grande Valley
cs.CV
Submitted: 2025-10-30
Updated: 2026-09-27
Importance score: 88/100
The gist: AD-SAM presents a fine-tuned vision foundation model designed for semantic segmentation in autonomous driving, significantly enhancing performance over existing models by integrating dual encoders
Key concepts
- Dual Encoder System
- This architecture uses two different neural network components working together. One, the SAM Vision Transformer (ViT-H), captures broad, global context of the entire scene. The second is a trainable ResNet-50 backbone that focuses on extracting fine, local spatial details necessary for precise road segmentation.
- Deformable Fusion Module
- This module intelligently combines features from the two encoders. It uses offset fields and modulation masks to adaptively align heterogeneous features across different scales and object geometries. This allows the model to effectively merge global context with local spatial information in a geometrically aware manner.
- Hybrid Loss Function
- The model is trained using four different loss terms simultaneously. These include Focal loss for class imbalance, Dice loss for region overlap optimization, LovászSoftmax for smooth IoU optimization, and Surface loss to improve boundary accuracy. This combination ensures the model learns robust segmentation features.
- Multi-Stage Decoder
- The decoder refines the fused multi-scale features through three sequential stages using deformable convolution. Each stage progressively reduces feature resolution (from 1024 down to 64 channels), allowing the model to iteratively refine its predictions from coarse semantic maps to fine pixel-level segmentation.
Terminology
Summary
AD-SAM presents a fine-tuned vision foundation model designed for semantic segmentation in autonomous driving, significantly enhancing performance over existing models by integrating dual encoders and a hybrid loss function. The core contribution lies in adapting the Segment Anything Model (SAM) to handle the spatial and geometric complexities of road scenes, leading to superior accuracy and improved data efficiency critical for reliable AD perception.
How it works
The AD-SAM architecture is built around a dual-encoder system that combines global context capture with local spatial detail extraction. The first encoder utilizes the Vision Transformer (ViT-H) from SAM, which processes inputs through multi-head self-attention to produce feature embeddings, while the second component employs a trainable convolutional deep learning backbone, specifically ResNet-50. This combination allows the model to capture both global semantic context
and local spatial detail.
Feature fusion is achieved through a deformable fusion module
that aligns these heterogeneous features across different scales and object geometries. This module computes offset fields and modulation masks to adaptively transform input features, defined by the equation:
y(p2) = 5w1/314/cdot x(p2 + p1 + Δp1) ⋅ m1 (Equation 1). Following fusion, a channel attention mechanism recalibrates the features through parallel average-pooling and maxpooling paths processed by a shared Multi-Layer Perceptron (MLP) to generate an attention vector.
How it works
The second core component is the multi-stage decoder,
which performs progressive multi-stage refinement using deformable attention. The decoder receives the concatenated multi-scale feature tensor, which is produced after the dual encoder's fusion step. This refinement occurs through three sequential stages:
-
Stage 1: DeformConv(1024→256) + GN(32) + GELU + Dropout(0.1)
-
Stage 2: DeformConv(256→128) + GN(16) + GELU + Dropout(0.1)
-
Stage 3: DeformConv(128→64) + GN(8) + GELU + Dropout(0.1)
The final class predictions are generated via a 3×3 deformable convolution projecting the 64-dimensional features to 19 semantic classes.
How it works
Training is guided by a hybrid loss
that integrates four complementary loss terms to improve performance across various aspects of segmentation:
((a) Cityscapes Dataset) Focal loss [26] with α = 0.25 and γ = 2 addresses class imbalance inherent in driving scenes.
((b) BDD100K Dataset) Dice loss [27] directly optimizes region overlap, measured using the Intersection of Union (IoU).
((c) BDD100K Dataset) LovászSoftmax loss [28] provides a smooth surrogate for discrete IoU optimization.
((d) BDD100K Dataset) Surface loss [29] enhances boundary accuracy by weighting errors using distance transforms.
How it works
The experimental setup involves evaluating AD-SAM on two benchmarks: Cityscapes and Berkeley DeepDrive 100K (BDD100K). Performance is measured using Intersection over Union (IoU) and mean Intersection over Union (mIoU). Training is performed on a single NVIDIA GeForce RTX 4090 GPU, utilizing mixed-precision training to accelerate the process. The model is trained for 100 epochs with a batch size of 2 per GPU, while the SAM's pretrained ViT-H encoder remains frozen.
How it works
The results demonstrate AD-SAM surpasses SAM, Generalized SAM (G-SAM), and DeepLabV3 in segmentation accuracy on both datasets. On Cityscapes, AD-SAM achieves 68.14 mIoU, while on BDD100K, it attains 59.50 mIoU. Furthermore, AD-SAM exhibits faster learning dynamics
and smoother convergence profile,
converging within 30-40 epochs and demonstrating strong data efficiency with only 100 samples achieving a competitive mIoU of 0.607 on Cityscapes, suggesting it is critical for reducing annotation costs in AD perception. Cross-domain retention analysis confirms a retention score of 0.8732, outperforming SAM (0.7622) and G-SAM (0.6754), indicating superior robustness under domain shift compared to DeepLabV3's 95.16 score, while offering greater adaptation flexibility
for real-world AD pipelines.
Improvements for AI systems
Based on the provided research paper detailing the AD-SAM model, here are specific improvements that can be made to existing AI systems, and what those improved systems could achieve:
-
Improve semantic segmentation accuracy in autonomous driving (AD) by adopting a dual-encoder architecture combining Vision Transformers (ViT) for global context and ResNet-50 for local spatial detail.
-
Enhance feature alignment across different scales and object geometries using a deformable fusion module that learns adaptive spatial transformations based on features from both encoders.
-
Improve the stability and accuracy of semantic segmentation models under class imbalance, boundary complexity, and noise by training with a hybrid loss function integrating Focal Loss, Dice Loss, Lovász-Softmax loss, and Surface Loss.
-
Increase data efficiency for AD perception tasks by fine-tuning foundation models (like SAM) with targeted architectural enhancements (dual encoder/deformable decoder) and optimized hybrid losses, allowing high performance even with limited labeled data (e.g., achieving competitive mIoU on Cityscapes with only 1,000 samples).
-
Improve cross-domain generalization in AD perception by fine-tuning foundation models on diverse datasets (like Cityscapes) and demonstrating robust transfer to novel environments (like BDD100K), achieving high retention scores (e.g., 87.32 retention score for AD-SAM vs. 76.22 for SAM).
-
Develop more reliable and stable learning dynamics in segmentation models by utilizing progressive multi-stage decoders with deformable attention, which enables precise boundary delineation crucial for complex urban scenes and improves convergence speed (e.g., converging within 30-40 epochs).
-
Increase the robustness of perception to domain shift (varying weather, lighting, road layouts) by leveraging foundation model pretraining and semi-supervised learning techniques that adapt to distribution shifts effectively without extensive manual relabeling.
-
Optimize computational efficiency for real-time AD inference by balancing high accuracy with moderate runtime growth, ensuring the system remains practical for deployment on embedded hardware (though further pruning or lightweight decoder designs are recommended).
-
Improve per-class segmentation precision, particularly for small or dynamic object classes (like poles, motorcycles), by incorporating class-aware sampling strategies or instance-level refinement modules to address intra-class variability and limited pixel footprints.
Sources
- Improving the Generalization of Segmentation Foundation Model under Distribution Shift via Weakly Supervised Adaptation
- Annotation Free Semantic Segmentation with Vision Foundation Models
- Segment Anything
- SAM 2: Segment Anything in Images and Videos
- Image Segmentation in Foundation Model Era: A Survey
- A Survey on Segment Anything Model (SAM): Vision Foundation Model Meets Prompt Engineering
- A Comprehensive Survey on Segment Anything Model for Vision and Beyond
- A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models
- A Review on Deep Learning Techniques Applied to Semantic Segmentation
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Deformable Convolutional Networks
- A Generalized Surface Loss for Reducing the Hausdorff Distance in Medical Imaging Segmentation
- Mixed Precision Training
- Rethinking Atrous Convolution for Semantic Image Segmentation
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models