RT-SEMamba: Real-Time Speech Enhancement Mamba via Progressive Knowledge Distillation
Academia Sinica · National Taiwan University · Kore University of Enna · University of Palermo · NVIDIA
cs.SD, cs.CL
Submitted: 2026-08-12
Updated: 2026-09-16
Comments: Accepted to INTERSPEECH 2026
Code: https://github.com/RoyChao19477/RT-SEMamba
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 75/100
The gist: RT-SEMamba is a fully causal speech enhancement (SE) model built upon causal time–frequency Mamba blocks.
Terminology
Summary
RT-SEMamba is a fully causal speech enhancement (SE) model built upon causal time–frequency Mamba blocks. Unlike Transformer-based architectures that rely on a growing key–value cache, Mamba propagates a fixed-size recurrent state per layer, enabling memory- and bandwidth-efficient long-form inference. The paper introduces a progressive knowledge distillation (KD) strategy that compresses an 8-layer teacher into a shallow 1-layer student by jointly distilling complex spectral outputs and intermediate representations. On Voicebank-DEMAND, the 8-layer RT-SEMamba achieves 3.32 PESQ with a 25 ms algorithmic latency constraint, and the distilled 1-layer student improves over a naive 1-layer baseline from 3.06 to 3.18 PESQ while preserving the same steady-state RTF, delivering a 2.75× speedup over the teacher. These results demonstrate that state-space models with progressive KD provide a competitive quality–latency trade-off for real-time SE.
The architecture operates in the complex STFT domain, taking magnitude and phase as input and predicting enhanced magnitude, phase, and complex reconstructed spectrum. It is made fully causal through several modifications: causal STFT/iSTFT at 16 kHz with window size 400 and hop 100, giving 25 ms algorithmic latency; asymmetric causal padding for all temporal convolutions; channel-wise LayerNorm with causal padding; and a uni-directional Time Mamba along time while frequency-axis modeling remains bidirectional. For streaming inference, the model operates in a 1-frame-in/1-frame-out mode, propagating a temporal frame buffer, a conv state buffer, and an ssm state, avoiding redundant recomputation and keeping per-frame compute independent of sequence length.
The knowledge distillation uses an 8-layer teacher and 1- or 2-layer students. The student is initialized with teacher weights for encoder, decoder, and the single cTF-Mamba block. Output-level distillation aligns magnitude, phase, and complex predictions with weights wmag=1.0, wpha=0.3, wcom=0.5. Intermediate feature distillation aggregates teacher block outputs by averaging across all 8 blocks, normalizes features per sample, and uses a feature distillation loss. A progressive ramp-up schedule gradually introduces the distillation signal over 10% of total training steps, with the final objective combining the original task loss plus weighted distillation losses (λout=0.5, λfeat=0.1).
Experiments on VCTK-DEMAND show that increasing cTF-Mamba depth from 1 to 8 layers improves PESQ from 3.06 to 3.32, with the largest gain from 1 to 2 layers (3.06→3.19) and diminishing returns beyond 4 layers. Streaming cost scales almost linearly with depth: parameters grow from 1.05M to 2.74M, MACs from 20.56 to 47.35 G/s, and RTF from 0.11 to 0.29. The 8-layer model is selected as teacher. For the 1-layer student, KD improves PESQ from 3.06 to 3.18; the 2-layer distilled model improves from 3.19 to 3.22 PESQ, approaching the 3-layer model's performance. Gap recovery for the 1-layer student is 46.2%, 19.2%, 63.6%, and 34.5% for PESQ, CSIG, CBAK, and COVL, respectively, with no additional streaming cost (RTF remains 0.11, about 2.6× faster than the 8-layer teacher).
Ablation studies explore hybrid Mamba–Transformer teachers. For 3-layer teachers, swapping the 2nd block with a cTF-Transformer yields virtually no gain (3.20 vs. 3.19 PESQ). For 5-layer teachers, replacing the 4th block improves PESQ from 3.27 to 3.32, matching the 8-layer all-Mamba model, but requires maintaining a KV cache in addition to the recurrent Mamba state, so the 8-layer all-Mamba model is adopted as teacher.
Comparison with prior causal and real-time models shows RT-SEMamba achieves competitive performance under strict latency. The 1-layer KD student reaches 3.18 PESQ with 1.05M parameters and 25 ms latency; the 2-layer KD student reaches 3.22 PESQ; the 8-layer model attains 3.32 PESQ at 2.74M parameters, all with the same 25 ms delay. The distilled models offer competitive PESQ and COVL while using compact architectures and strictly bounded latency, making them attractive for practical real-time enhancement under tight latency and data constraints.
Improvements for AI systems
Improvements to AI systems:
-
Efficient long-form inference for sequence models: Replace Transformer-based KV caches with fixed-size recurrent state-space layers (Mamba) to enable memory- and bandwidth-efficient streaming inference. The improved system can process arbitrarily long audio streams with constant per-frame compute and memory, independent of sequence length.
-
Progressive knowledge distillation for model compression: Implement a two-stage distillation (output-level + intermediate feature-level) with a ramp-up schedule to compress deep teachers into shallow students. The improved system can recover 46.2% of PESQ gap while achieving 2.75× speedup, enabling deployment on low-resource devices without retraining from scratch.
-
Causal audio processing pipeline: Adopt fully causal STFT/iSTFT with asymmetric padding and uni-directional temporal modeling. The improved system can perform real-time speech enhancement with strict 25 ms algorithmic latency, suitable for hearing aids, teleconferencing, or voice assistants.
-
Hybrid architecture selection via ablation: Use systematic ablation (e.g., swapping Mamba blocks with Transformer) to decide when recurrent state is sufficient vs. when attention adds value. The improved system can automatically choose the most latency-efficient architecture that meets quality targets, avoiding unnecessary KV cache overhead.
-
Teacher-student weight initialization for KD: Initialize student encoder/decoder and core blocks from teacher weights to accelerate convergence. The improved system can train compact models faster and with less data, achieving near-teacher quality (3.18 vs. 3.32 PESQ) with 62% fewer parameters.
-
Feature normalization in distillation: Normalize intermediate features per sample before computing distillation loss. The improved system can stabilize training when distilling from multi-block teachers, preventing scale mismatches and improving gradient flow.
What the improved AI system can do:
-
Perform real-time speech enhancement on live audio streams (e.g., in hearing aids or mobile calls) with only 25 ms latency, using a 1.05M-parameter model that runs 2.75× faster than a larger teacher while maintaining 3.18 PESQ.
-
Deploy on edge devices with limited memory and bandwidth, as it requires only a fixed-size recurrent state per layer instead of growing caches.
-
Be trained efficiently from a large teacher using progressive KD, recovering most of the quality gap (46.2% for PESQ) without extra inference cost.
-
Adapt to strict latency constraints (e.g., 25 ms) while maintaining competitive quality (3.22 PESQ with 2-layer student), making it viable for interactive applications.
-
Scale depth flexibly (1–8 layers) to trade quality vs. compute, with predictable RTF scaling (0.11 to 0.29) and no hidden memory growth.
Sources
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces
- SPMamba: State-space model is all you need in speech separation
- Distilling the Knowledge in a Neural Network
- Jamba: A Hybrid Transformer-Mamba Language Model
Related papers
- Few-Shot Open-Set Audio Classification via Transductive Prototype Refinement and Class Logit Enhancement
- Spectral Masking and Interpolation Attack (SMIA): A Black-box Adversarial Attack against Voice Authentication and Anti-Spoofing Systems
- AVMeme Exam: A Multimodal Multilingual Multicultural Benchmark for LLMs' Contextual and Cultural Knowledge and Thinking
- SoundWeaver: Compositional Warm-Starting for Text-to-Audio Diffusion Serving
- WASIL: In-the-Wild Arabic Spoken Interactions with LLMs
- Efficient Audiovisual Speech Processing via MUTUD: Multimodal Training and Unimodal Deployment