RetiWave-Mamba: A Dual-Stream Network for Retinal Disease Detection based on Multi-scale Context and Feature-Adaptive Mamba Projection
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "RetiWave-Mamba: A Dual-Stream Network for Retinal Disease Detection based on Multi-scale Context and Feature-Adaptive Mamba Projection".
Jane: Retinal diseases pose a significant global health challenge, requiring early and accurate diagnosis, which necessitates automated analysis of Optical Coherence Tomography (OCT) images to overcome manual interpretation difficulties.
Tom: First, who's behind it and why it matters.
Title and authors: Tom: So, we're starting with the title and authors for this paper, 'RetiWave-Mamba: A Dual-Stream Network for Retinal Disease Detection based on Multi-scale Context and Feature-Adaptive Mamba Projection.' It sounds pretty technical, but it points to a really specific solution for a big problem in medical imaging.
Jane: I'm curious about what that actually means for someone looking at an OCT scan. Essentially, they’re saying this framework uses two different ways to look at the image—one focusing on the overall structure and another focusing on the tiny details and frequencies—to figure out if there's a retinal disease.
Lu: From a research standpoint, it’s interesting that they are explicitly combining spatial domain analysis with frequency domain learning, which is a sophisticated way to tackle noise when you're dealing with complex biological structures like the retina.
Meng: From an engineering side, I wonder if this dual-stream approach makes the model more stable when we feed it noisy data from real scanners; stability is always a big concern for deployment.
Lalam: I think what's really exciting here is how this architecture could fundamentally shift how we train diagnostic models to be more resilient to the kind of noise that plagues medical scans.
The paper's summary: Tom: Okay, so summarizing the main idea of 'RetiWave-Mamba: A Dual-Stream Network for Retinal Disease Detection based on Multi-scale Context and Feature-Adaptive Mamba Projection,' they’re proposing a framework that breaks down OCT images using a Discrete Wavelet Transform into low and high frequency streams.
Jane: That decomposition lets them process the general structure in one stream, which is then enhanced by a Multi-scale Contextual Localization Module, while the other stream handles the fine details using an Attention-Guided High-Resolution Network and a Frequency-Adaptive Mamba Projector.
Lu: The paper highlights that this dual design lets the system capture both global structural contexts and detailed frequency features simultaneously, which they argue is key to improving performance in noisy settings.
Meng: From a practical standpoint, decoupling the processing like that suggests that if one part of the image processing struggles with noise, the other stream might still provide a reliable signal for classification.
Lalam: I think this approach to separating global context from fine texture is really powerful because it addresses the challenge of having lesions of different scales in a single image.
The paper's improvements: Tom: Moving on to what they actually improved, the authors put forward several specific enhancements for this framework. They focused on building that low-frequency stream with a Multi-scale Contextual Localization Module, which uses multi-scale dilation fusion and spatial attention to sharpen lesion localization.
Jane: And in the high-frequency branch, they introduced an Attention-Guided Fusion mechanism that selectively filters information during multi-scale interactions to suppress noise propagation effectively.
Lu: The Frequency-Adaptive Mamba Projector is another major contribution; it’s designed to handle long-range dependencies within the fine textural details, which is crucial for telling similar lesions apart.
Meng: I'm thinking about the practical benefit of that FAMP module—if it can model those long-range textures better, it means we might see higher accuracy when distinguishing between subtly different types of retinal issues in real clinical scenarios.
Lalam: For me, the most impactful improvement is definitely how they use the Mamba projection to capture those long-range dependencies; that could lead to much more reliable AI systems when diagnosing complex conditions.
Conclusion: Tom: So, wrapping up our discussion on 'RetiWave-Mamba: A Dual-Stream Network for Retinal Disease Detection based on Multi-scale Context and Frequency-Adaptive Mamba Projection,' the authors show how combining spatial and frequency domain learning can lead to a very robust system.
Jane: They managed to achieve a state of the art classification accuracy of ninety-eight point two five percent on the OCT-C8 dataset by using this dual stream architecture, which is impressive considering the challenges they were tackling upfront.
Lu: The combination of MCLM, AG-HRNet, and FAMP is a really creative way to leverage different feature representations to get a comprehensive understanding of both structure and texture in one model.
Meng: From an engineering view, the result is that we have a very strong classifier that handles noisy inputs well, which means it could be deployable sooner than if we relied on models that are overly sensitive to image quality fluctuations.
Lalam: I feel really optimistic about this work; the implications for AI in healthcare are huge because it shows how deep learning can systematically tackle the noise and complexity inherent in medical data to achieve high precision.
Tom: Exactly, so 'RetiWave-Mamba' gives us a solid blueprint for building more resilient diagnostic tools that don't just look at one aspect of the image but understand both structure and fine detail.
Jane: It’s a powerful tool because it moves beyond just relying on standard CNNs by integrating these sophisticated modules to handle the complexity of retinal pathology.
Lu: This paper really sets a good direction for how we can use state space models alongside traditional image processing techniques to model complex visual data effectively.
Meng: We need to keep an eye on how this architecture performs when we move it off the lab bench and onto actual clinical hardware where the image acquisition conditions aren't perfectly controlled.
Lalam: It’s exciting because this level of detail in model design shows that we can engineer AI systems that are not just accurate in a perfect setting, but actually robust enough for messy, real-world medical applications.
Cheng Cheng, Jin Hong
School of Information Engineering, Nanchang University
cs.CV
Submitted: 2026-08-18
Updated: 2026-10-02
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 90/100
The gist: Retinal diseases pose a significant global health challenge, requiring early and accurate diagnosis, which necessitates automated analysis of Optical Coherence Tomography (OCT) images to overcome
Key concepts
- Discrete Wavelet Transform (DWT)
- This technique breaks down the input OCT image into different frequency components: one low-frequency component representing broad structure and three high-frequency components capturing fine details. This decomposition allows the model to analyze both global context and specific textural information simultaneously for better diagnosis.
- Low-Frequency Stream
- This branch focuses on preserving the overall structural context of the image while naturally filtering out high-frequency noise. It uses a backbone enhanced by a Multi-scale Contextual Localization Module (MCLM) to accurately pinpoint where lesions are located based on wide contextual information.
- Mamba Projection
- Mamba is a specialized State Space Model used to efficiently capture long-range dependencies in the high-frequency stream. It helps the network understand complex, fine-grained textural details and edge information that are crucial for distinguishing between similar retinal pathologies.
Terminology
Summary
Retinal diseases pose a significant global health challenge, requiring early and accurate diagnosis, which necessitates automated analysis of Optical Coherence Tomography (OCT) images to overcome manual interpretation difficulties. The proposed framework integrates spatial-frequency domain learning with State Space Models to robustly identify retinal pathologies under noisy conditions.
The gist: The RetiWave-Mamba framework achieves a state-of-the-art classification accuracy of 98.25% on the OCT-C8 dataset by combining a dual-stream architecture with Discrete Wavelet Transform decomposition, spatial attention modules, and Mamba projection for capturing both global structural context and fine-grained frequency details.
Framework Overview
The RetiWave-Mamba framework is a novel hybrid dual-stream network designed to address the challenges of speckle noise and varying lesion scales in OCT image analysis. The core architecture systematically processes global structural context and fine-grained frequency details in parallel for robust retinal disease classification. This design leverages the strengths of different feature representations, significantly improving the model's performance in noisy environments.
The overall framework is structured around three key components:
-
A Discrete Wavelet Transform (DWT) block that decomposes the input image into one low-frequency (LF) component and three high-frequency (HF) components.
-
A low-frequency stream equipped with a Multi-scale Contextual Localization Module (MCLM).
-
A high-frequency stream integrating an Attention-Guided High-Resolution Network (AGHRNet) with a Frequency-Adaptive Mamba Projector (FAMP).
Low-Frequency Stream: Spatial Context Modeling
The low-frequency branch preserves the global structure and contextual information while inherently suppressing high-frequency noise. This stream utilizes a ResNet backbone structured into four stages of Bottleneck Blocks, enhanced by the Multi-scale Contextual Localization Module (MCLM) inserted after the first three stages. The MCLM is designed to enhance the global receptive field and precisely localize lesion regions by synergizing multi-scale dilation with spatial attention.
The MDF unit within MCLM consists of two cascaded sub-units:
-
The Multi-scale Dilation Fusion (MDF) unit, which captures wide-range contextual information using a projection function followed by element-wise summation of features from a standard convolution branch and three dilated convolution branches with rates = 6, 12, and 18.
-
The Multi-Scale Wavelet Spatial Attention (MSW SA) unit, which refines the features to localize lesion regions by generating a refined spatial attention map (MS) via aggregation of channel information and applying it to the intermediate feature map F via element-wise multiplication: y = F ⊗ MS.
High-Frequency Stream: Fine-Grained Detail Extraction
The high-frequency branch is initiated by concatenating the three HF components (LH, HL, and HH sub-bands) into a 9-channel input. This stream is handled by the Attention-Guided High-Resolution Network (AGHRNet), which maintains four parallel high-to-low resolution branches to preserve feature map fidelity.
The core innovation here is the Attention-Guided Fusion (AGF) mechanism, which operates as an exclusive pathway for all cross-resolution feature exchanges. This mechanism generates an attention gate G from the source features and applies it to the incoming features via element-wise multiplication: yfused = Ph + (Pl' ⊗ G). This gate operation acts as a selective filter, re-weighting incoming features to amplify useful information and suppress noise propagation during multi-scale interactions.
Frequency-Adaptive Mamba Projection (FAMP)
The Frequency-Adaptive Mamba Projector (FAMP) is the specialized module designed to process the four multi-scale feature maps output by the AGHRNet. Its primary function is to capture long-range dependencies within the fine-grained textural and edge details critical for distinguishing between morphologically similar lesions.
FAMP operates through a frequency-gating mechanism:
-
A channel-wise attention mask (M) is generated via global average pooling and a 1×1 convolution, which splits the input feature map X into two distinct paths: XL = X ⊗ M and XH = X ⊗ (1 − M).
-
The components are concatenated along the channel dimension, flattened, and normalized to form a sequence representation (Xs).
-
This sequence is fed into the Mamba module [34] to efficiently model global dependencies, yielding a sequence output Xs seq.
-
A weight generator maps this back to the channel dimension via a sigmoid function to produce adaptive weights (WL and WH), which are then used to modulate the original frequency paths before being combined via a Fusion block (Ffusion).
Performance and Robustness
The model achieved a state-of-the-art (SOTA) accuracy of 98.
Improvements for AI systems
Based on the provided scientific paper, here are the specific improvements that can be made to existing AI systems and what those improved systems will be capable of:
) Improvements for Existing AI Systems:
-
[RetiWave-Mamba Dual-Stream Framework Integration]: Implement a dual-stream architecture that decouples structural context (low-frequency stream) from fine-grained textural details (high-frequency stream).
-
[Frequency Decomposition via DWT]: Use the Discrete Wavelet Transform (DWT) to decompose input OCT images into low-frequency structural approximations and high-frequency detail components.
-
[Enhanced Low-Frequency Contextual Localization]: Integrate a Multi-scale Contextual Localization Module (MCLM) in the low-frequency branch, combining multi-scale dilation fusion with spatial attention to expand the global receptive field and precisely localize lesion regions without losing resolution.
-
[Noise Suppression in High-Frequency Stream]: Implement an Attention-Guided High-Resolution Network (AG-HRNet) equipped with an intelligent gating mechanism (Attention Gated Fusion, AGF) during multi-scale interactions to selectively amplify diagnostic textures while suppressing background speckle noise propagation.
-
[Long-Range Dependency Modeling for Textures]: Incorporate a Frequency-Adaptive Mamba Projector (FAMP) in the high-frequency branch. This module uses the selective scan mechanism of State Space Models (Mamba) to capture long-range dependencies across disjoint high-frequency textural features, enabling robust distinction between morphologically similar lesions (e.g., CNV vs. Drusen).
-
[Robustness via Progressive Noise Injection]: Integrate a progressive noise injection training curriculum (Phases 1–3) into the training pipeline to force the model to learn invariant pathological features even under high-intensity speckle noise, ensuring clinical readiness in real-world, low-quality scans.
) Capabilities of the Improved AI System:
The improved AI system will be a highly robust and clinically valuable automated diagnostic tool for retinal diseases, capable of:
-
[SOTA Classification Accuracy]: Achieve state-of-the-art classification accuracy (up to 98.25% on OCT-C8) by capturing both global structural context and fine textural details simultaneously.
-
[Precise Lesion Localization]: Accurately delineate lesion boundaries with high precision, overcoming the limitations of standard CNNs in balancing global context understanding with precise localization via the MCLM module.
-
[Superior Inter-Class Discrimination]: Resolve critical diagnostic ambiguities between morphologically similar retinal pathologies (like CNV and Drusen) by leveraging frequency-domain analysis and long-range dependency modeling (FAMP), which filters out noise that typically blurs these subtle textural differences.
-
[Clinical Robustness in Noisy Environments]: Maintain high diagnostic accuracy (e.g., >97.79% under severe noise conditions) by explicitly decoupling signal from speckle noise through frequency-domain learning and an attention-guided gating mechanism, making it reliable for deployment on real-world clinical scanners where image quality is variable.
-
[Efficient Deployment]: Offer a computationally efficient alternative to complex Transformer models (like Swin Tiny), achieving superior performance with significantly lower parameter counts and GFLOPs compared to methods like ASPP, making it suitable for rapid inference in resource-constrained clinical settings.
Sources
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Multi-Scale Context Aggregation by Dilated Convolutions
- BAM: Bottleneck Attention Module
- Efficiently Modeling Long Sequences with Structured State Spaces
- Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models