MambaDSF: Multi-Scale SSM with Dilated Feature Fusion for Sonar Small Target Detection

arXiv:2605.24928 · cs.CV · Submitted 2026-05-24 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "MambaDSF: Multi-Scale SSM with Dilated Feature Fusion for Sonar Small Target Detection".

Jane: Sonar imaging presents significant challenges for detecting small targets due to insufficient pixel coverage, low acoustic contrast, and scale ambiguity across imaging ranges.

Tom: First, who's behind it and why it matters.

Paper summary: Tom: So Jane and I have been looking at this paper titled "MambaDSF: Multi-Scale SSM with Dilated Feature Fusion for Sonar Small Target Detection," and it's really interesting because sonar imaging has these real issues with small targets due to low contrast and scale ambiguity.

Jane: Yeah, Tom, the main thesis seems to be that they've developed a hybrid framework that combines a MambaEFP backbone with a Dilate Fusion Mamba encoder and specific task-dependent losses to tackle those challenges.

Lu: I think what really catches my eye is how they manage the linear complexity while still capturing both local echo cues and global acoustic context, which is exactly what you need for this kind of detection task.

Meng: From an engineering standpoint, that linear cost aspect with Mamba models is appealing because we're dealing with potentially massive datasets in sonar imaging where efficiency matters a lot.

Lalam: I see the potential here to improve how we process underwater data; imagine this framework applied across different sensor modalities, it could really enhance our cultural understanding of remote sensing.

Tom: Exactly, and what they claim is that MambaDSF achieves ninety-one point five percent mAP50 on the UATD forward-looking sonar benchmark with twenty-eight point seven million parameters, which they say surpasses all the detectors they compared it against.

Jane: That performance figure is substantial, and it really shows how effective their design is at handling those small targets that are usually so hard to find in sonar data.

Lu: The MambaEFP backbone itself has some clever components like the MambaVision feature extractor which uses selective SSM blocks where the input-dependent parameter Delta controls whether token information is retained or noise is suppressed, which sounds very intuitive for clutter attenuation.

Tom: That's a key part of their methodology, and then you have that Hybrid Block addressing one-dimensional flattening by using a dual-branch design with depthwise convolutions to preserve those fine echo boundaries without sequence flattening.

Meng: Preserving fine-grained echo boundaries without sequence flattening is important because if you lose that detail, the detection accuracy for small targets drops immediately.

Jane: And then they layer on this Dilate Fusion Mamba encoder, which seems to be the mechanism for aligning multi-scale features across different pyramid levels and giving it multi-receptive-field sensitivity.

Lalam: The way they fuse information across scales sounds like a very robust way to ensure that a small target's signal isn't just missed because the network was looking at the wrong resolution.

Tom: They achieve this refinement through two steps: first, intra-scale enhancement using Multi-Scale Dilated Attention to extract local cues, and then cross-scale fusion using Selective Cross-scale Modulation and Aggregated Feature Refinement.

Lu: The SCM Block uses FusSSM whose parameters are derived from auxiliary features of adjacent pyramid levels, which is a sophisticated way to derive the gating parameters based on spatial alignment.

Meng: So it's not just looking at different scales in isolation; they are actively using information from neighboring scales to modulate how the SSM operates, which makes sense for context.

Jane: And finally, they incorporate task-specific losses like Scale-Adaptive Weighted IoU and Cross-Scale Coherence to stabilize training specifically for small targets.

Lalam: Those losses are smart; scaling the IoU loss with a Gaussian model and using the Wasserstein distance seems tailored precisely to give more weight to those tiny ground truth boxes during training.

Tom: That’s how they tackle the gradient vanishing issue, and then there's the Cross-Scale Coherence loss that penalizes inconsistent multi-scale representations across different encoder levels.

Lu: The results are really impressive because they show this synergy; when you look at the ablation study, Component B with the full DFMamba encoder gave them the largest single-factor mAP50 gain, confirming that multi-receptive-field coverage is a dominant factor for handling diverse sonar echo sizes.

Jane: It sounds like MambaDSF really nails the architectural requirements for this specific problem by combining efficient linear complexity with rich multi-scale context modeling.

Meng: From an engineering standpoint, achieving forty-seven point five frames per second with twenty-eight point seven million parameters is a very favorable trade-off when you have to balance accuracy and processing speed for real-time sonar applications.

Lalam: The implications for culture could be huge; if this kind of efficient, context-aware detection can be developed, it could open up new possibilities in monitoring complex underwater environments globally.

Tom: So, to wrap up this discussion on "MambaDSF: Multi-Scale SSM with Dilated Feature Fusion for Sonar Small Target Detection," we've seen how they tackle scale ambiguity and noise suppression using selective SSM blocks and multi-scale fusion techniques.

Jane: We've also discussed how the task-specific losses help stabilize training for those tiny targets, and the overall performance metrics show strong results on the UATD benchmark.

Lu: The core contribution seems to be demonstrating that linear complexity models can incorporate complex multi-scale dependencies effectively through their fusion architecture.

Meng: What I find most practical is how they manage the computational cost while maintaining high accuracy, especially when you consider the need for speed in real-world deployment.

Lalam: The broader impact is about making AI tools capable of understanding and accurately perceiving nuanced, low-contrast data from complex physical environments across many different domains.

Conclusion: Tom: So we've been digging into MambaDSF, and now it's time to wrap up what this paper really is all about—the title and the authors are a great place to start.

Jane: Yeah, Tom, I think understanding the name "Multi-Scale SSM with Dilated Feature Fusion for Sonar Small Target Detection" helps us get a feel for exactly what they tackled.

Lu: From my side, I find it fascinating how they combine the Mamba architecture with explicit multi-scale fusion techniques; it’s a really creative way to handle the scale ambiguity inherent in sonar data.

Meng: I'm thinking about the authors too; seeing how they balanced linear complexity with high performance is something that makes me wonder about its practical implementation down the line.

Lalam: I think this paper points toward a future where AI can perceive and interpret incredibly complex, low-contrast physical environments in ways we haven't even imagined yet.

Tom: That’s right, Lalam; the implication here is that we can start to build AI systems that are much better at finding those tiny echoes in noisy sonar scans than before.

Jane: In simple terms, the paper explains how they used a specific neural network structure to look at different parts of an image at different resolutions simultaneously for better results.

Lu: Exactly, Jane; the Mamba component lets them model long-range context efficiently while the Dilated Fusion Mamba encoder makes sure they see both fine details and broad surroundings together.

Meng: The practical impact I'm seeing is that this method could lead to faster, more reliable sonar systems in underwater exploration or even for non-destructive testing where detecting small defects matters.

Lalam: And on a broader cultural level, imagine the kind of environmental monitoring we could do with these improved detection capabilities; it opens up new frontiers in understanding our planet's subsurface structure.

Tom: So, to summarize this conclusion, MambaDSF is essentially a sophisticated way to give AI the vision needed to reliably spot small targets in challenging sonar images by smartly fusing information across multiple scales.

Jane: That’s a solid way to put it; they are using a clever combination of architecture and loss functions so that the AI doesn't miss those tiny echoes.

Lu: I think the real takeaway is showing that linear complexity models, like Mamba, can still handle these kinds of intricate multi-scale relationships very effectively when you design the fusion layers just right.

Meng: And for engineering teams, it means we have a proven blueprint for how to structure these large models to ensure they maintain high accuracy without requiring impossible amounts of computational power.

Lalam: This work shows us that the next wave of AI advancement won't just be about bigger models, but about designing architectures that are intelligently fused and tuned for specific, difficult sensory tasks.

Tom: Absolutely; this paper sets a new standard for how we approach small target detection in challenging sensing domains.

Jane: And it makes me wonder what other types of sensor data might benefit from this kind of multi-scale fusion approach to improve their performance.

School of Information Science and Engineering, Ocean University of China

cs.CV

Submitted: 2026-05-24

Updated: 2026-10-07

Code: https://github.com/IDontKnowAAA/MambaDSF

Importance score: 89/100

The gist: Sonar imaging presents significant challenges for detecting small targets due to insufficient pixel coverage, low acoustic contrast, and scale ambiguity across imaging ranges.

Key concepts

MambaEFP Backbone
This backbone jointly captures local echo cues and global acoustic context using a Mamba architecture. It uses selective SSM blocks where input-dependent parameters filter noise, while a dual-branch design preserves fine echo boundaries without sequence flattening.
Dilate Fusion Mamba Encoder (DFMamba)
This encoder refines features by aligning information across different pyramid levels. It uses Multi-Scale Dilated Attention to encode local cues at various scales and Selective Cross-scale Modulation to fuse complementary evidence from adjacent pyramid levels, ensuring multi-receptive field sensitivity.
Scale-Adaptive Weighted IoU (SA-WIoU)
This loss function stabilizes training for small targets by using the Normalized Wasserstein Distance instead of standard IoU. It uses an area-adaptive weight to favor the Wasserstein distance when dealing with smaller ground truth boxes, improving localization accuracy.
Cross-Scale Coherence (CSC)
This loss penalizes inconsistent feature representations across different pyramid levels. It aligns feature vectors sampled at a target's center across these levels using pairwise cosine similarity, forcing the model to produce coherent multi-scale features.

Terminology

Summary

Sonar imaging presents significant challenges for detecting small targets due to insufficient pixel coverage, low acoustic contrast, and scale ambiguity across imaging ranges. This work proposes MambaDSF, a hybrid framework combining a MambaEFP backbone with a Dilate Fusion Mamba encoder and task-specific losses to achieve state-of-the-art performance in sonar small target detection.

The gist

MambaDSF achieves 91.5% mAP50 on the UATD forward-looking sonar benchmark with 28.7 M parameters, surpassing all compared detectors.

MambaEFP Backbone

The MambaEFP backbone is designed to jointly capture local echo cues and global acoustic context at linear complexity. It integrates three sub-components: a MambaVision feature extractor, a Hybrid Block, and an enhanced feature pyramid (E-FPN). The MambaVision feature extractor utilizes selective SSM blocks where the input-dependent parameter ∆ is central to sonar processing: large values retain target token information in the hidden state while small values suppress noise token contributions, providing scene-wide clutter attenuation at linear cost. Furthermore, the Hybrid Block addresses 1-D flattening by using a "dual-branch design: a local branch with parallel DW3×3 and DW5×5 depthwise convolutions operates directly in the spatial domain, preserving fine-grained echo boundaries without sequence flattening; its output is combined with the SSM branch through learnable residual weights." The E-FPN cascades three modules within each pyramid block: a contrast enhancement block, an edge attention block, and a multi-scale enhancer.

Dilate Fusion Mamba Encoder (DFMamba)

The DFMamba encoder refines the features by enforcing multi-scale feature alignment across pyramid levels and providing multi-receptive-field sensitivity. This is achieved through two steps:

  1. Intra-scale enhancement via Multi-Scale Dilated Attention (MSDA): At each level Ni, query, key, and value tensors are split into nd groups of C/nd channels, each processed at a distinct dilation rate d: Unfoldd extracts a 3×3 neighbourhood at pixel spacing d. This allows for joint encoding of local target cues and surrounding context.

  2. Cross-scale fusion via Selective Cross-scale Modulation (SCM) and Aggregated Feature Refinement (AFR): Each MSDA branch undergoes independent cross-scale fusion in the SCM Block via FusSSM, whose parameters derive from spatially aligned auxiliary features of adjacent pyramid levels. The modulator parameters are derived using equations involving a learnable projection Wp: ∆fus, Bfus, Cfus = SplitWp Fd m. This process ensures that the SSM parameters are derived from the modulator rather than the target branch, encoding cross-scale complementary evidence. Finally, the AFR module aggregates all branches and produces features for the decoder.

Task-Specific Losses

To stabilize training for small targets, two task-specific losses are devised:

  1. Scale-Adaptive Weighted IoU (SA-WIoU): This loss counters gradient vanishing by modeling each bounding box as a 2-D Gaussian and using the squared Wasserstein distance, yielding the Normalized Wasserstein Distance (LNWD). It blends this with CIoU via an area-adaptive weight: LSA-WIoU = ω LNWD + (1−ω)LCIoU, ω = exp (−ag / τs, where ag is the normalized ground-truth box area, favoring the Wasserstein term for small targets.

  2. Cross-Scale Coherence (CSC): This loss penalizes inconsistent multi-scale representations across encoder levels by aligning feature vectors sampled at each ground-truth box centre across pyramid levels in direction via pairwise cosine similarity: LCSC = 1 − 1 / PG Σ g∈G Σ (i,j)∈P e g i · e g j e g i e g j.

Experimental Results and Generalization

MambaDSF achieves the highest mAP50 of 91.5% on the full UATD test set and maintains high accuracy on small-target subsets (+2.2 pp gain). Cross-domain evaluation on FLS and MD-FLS confirms generalization, as MambaEFP’s global context suppresses domain-varying noise, DFMamba aligns multi-scale semantics across unseen target distributions, and SA-WIoU sharpens localization. The model achieves 47.5 FPS with 28.7 M parameters, demonstrating a favorable computational trade-off for the achieved accuracy.

Ablation Study Insights

The full factorial ablation study confirmed that the three components address distinct aspects synergistically: Component B (full DFMamba encoder) yielded the largest single-factor mAP50 gain, confirming "multi-receptive-field coverage is the dominant factor for handling diverse sonar echo sizes.

Improvements for AI systems

As a fastidious and diligent AI researcher, I have analyzed the MambaDSF (Multi-Scale State-Space Model with Dilated Feature Fusion for Sonar Small Target Detection) paper. The core innovation lies in creating a hybrid architecture that synergistically combines linear complexity (Mamba SSM) for global context modeling, multi-scale feature fusion (Dilated Attention/E-FPN), and specialized training losses to overcome the inherent challenges of small target detection in sonar imagery.

Here are the specific improvements and capabilities this framework enables:


)

  1. Acknowledge that current CNNs fail due to lack of global context, while Transformers fail due to quadratic computational cost, and Mamba-based models lack multi-scale alignment and small-target supervision.

  2. Improve AI systems by implementing the MambaDSF framework, which consists of three coordinated components:

  3. Design the MambaEFP backbone: This component jointly captures local echo cues (via a Hybrid Block preserving spatial detail) and global acoustic context (via MambaVision), enhanced by an E-FPN that explicitly applies contrast enhancement, edge attention, and multi-scale ASPP convolutions to amplify target echo intensity and boundary sharpness.

  4. Introduce the DFMamba encoder: This module enforces multi-scale semantic alignment across pyramid levels by splitting features into dilated attention branches (MSDA) and applying a Selective Cross-scale Modulation (SCM) block with FusSSM. This allows the model to handle targets whose apparent size varies significantly with imaging range by leveraging complementary receptive fields.

  5. Incorporate task-specific losses: Utilize the Scale-Adaptive Weighted IoU (SA-WIoU) loss, which uses a 2-D Gaussian model and Wasserstein distance weighted by object area, to ensure non-zero regression gradients for small targets even when overlap is zero. Additionally, implement the Cross-Scale Coherence (CSC) loss to penalize inconsistent multi-scale representations of the same target across different pyramid levels.

  6. Enable the improved AI system to perform state-of-the-art detection in challenging sonar environments, specifically:

  7. Detect small targets with high precision (achieving up to 91.5% mAP50 on UATD) by effectively suppressing noise-induced false alarms through global acoustic context modeling (Mamba's selective gating).

  8. Robustly handle scale ambiguity by maintaining semantic consistency across different imaging ranges, ensuring accurate localization regardless of the target's apparent size.

  9. Generalize detection capabilities across different sonar modalities (FLS and MD-FLS) due to the cross-domain evaluation validation, allowing the system to perform reliably on unseen acoustic data distributions.

  10. Achieve superior efficiency compared to Transformer baselines, operating at 47.5 FPS with only 28.7M parameters, making it suitable for real-time deployment on resource-constrained platforms like AUVs for autonomous underwater navigation and surveillance tasks.

Sources

Related papers