HybridSB-MoE: Dual-Domain Schr"odinger Bridges with Scene-Adaptive Expert Routing for Speech Enhancement

arXiv:2608.12715 · cs.SD, cs.AI · Submitted 2026-08-13 · Read on arXiv

Zhengyi Lu, Aswini Sivakumar, Jie Hu, Yao Qiang

Oakland University

cs.SD, cs.AI

Submitted: 2026-08-13

Updated: 2026-08-14

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 75/100

The gist: HybridSB-MoE is a dual-domain framework for speech enhancement that combines a heterogeneous spectral Mixture-of-Experts (MoE) pathway with a waveform Schrödinger Bridge (SB) pathway, unified by an

Terminology

Summary

HybridSB-MoE is a dual-domain framework for speech enhancement that combines a heterogeneous spectral Mixture-of-Experts (MoE) pathway with a waveform Schrödinger Bridge (SB) pathway, unified by an asymmetric uncertainty fusion design. The paper identifies three structural limitations in existing generative speech enhancement: (i) single-domain commitment, where methods operate in either the waveform or spectral domain, sacrificing the complementary inductive bias of the other; (ii) uniform processing across heterogeneous noise, where a single network handles stationary appliance hum, harmonic engine noise, and non-stationary crowd babble alike; and (iii) loosely controlled sampling cost, where generative pipelines require many iterative refinement steps with limited formal connection between training objective and inference budget.

The proposed framework addresses these gaps through three main contributions. First, asymmetric uncertainty fusion: the spectral path captures epistemic uncertainty via expert disagreement, while the waveform bridge models aleatoric variance through stochastic dynamics. These are fused asymmetrically, allowing the mixing weight to adapt to distinct error regimes rather than average predictions. Second, a heterogeneous MoE with top-k=2 routing across five distinct architectural archetypes (Home, Nature, Office, Transport, Public), where architectural diversity makes the epistemic signal indicate which inductive bias fails rather than small perturbations among similar experts. Third, a discretization bound (Theorem 1) showing that path-consistency and trajectory regularizers together bound the K-step bridge sampling error in 2-Wasserstein distance at rate K(-α), making small-K inference an objective-level guarantee rather than an empirical claim.

The spectral pathway applies the MoE to log-magnitude STFT features z = logS y, with two-level gating combining archetype-level routing (pooling across time) and token-level routing (frame-level refinement). The routed output is x̂spec = Σ i∈Ik G i(z) E i(z), where Ik indexes selected experts. The waveform pathway uses a Schrödinger Bridge formulation where the intermediate state interpolates between noisy observation y and clean target x: x t = √(β̄ t)x + √(1-β̄ t)y + σ t ε, with a cosine schedule and σ t = σ max√(β̄ t(1-β̄ t)) that vanishes at both endpoints. The reverse process uses a data-prediction update: x t-1 = √(β̄ t-1)x̂ θ(x t, y, t) + √(1-β̄ t-1)y + σ t-1z.

The path-consistency loss L path = E t,t'x̂ θ(x t, y, t) - x̂ θ(x t', y, t')2 enforces cross-timestep agreement of clean-signal predictions along the same trajectory. The trajectory loss L traj = E tx t - (√(β̄ t)x̂ θ(x t, y, t) + √(1-β̄ t)y)2 anchors x t to the schedule-consistent reconstruction. Theorem 1 states: W2(p̂ K, p br 0) ≤ C1K(-α) + C2√(L* path + L* traj), with α = min(1, γ), where L* are training-termination values. This shows that once the regularizers are minimized, the C1K(-α) term saturates at modest K, justifying the K=8 inference budget.

The asymmetric fusion combines the two pathways: x̂(t) = w·x spec(t) + (1-w)·x wave(t), where w = σ(MLP(ũ epi, ũ ale)) ∈ [0,1]. The epistemic uncertainty u epi = (1/kT f)Σ i∈IkE i(z) - Ē(z)22 measures expert disagreement, while the aleatoric uncertainty u ale comes from the U-Net's per-sample log-variance head. A calibration loss L cal = (u epi - x spec - x22)2 + (u ale - x wave - x22)2 anchors these scalars to actual reconstruction errors.

On VoiceBank+DEMAND, HybridSB-MoE achieves the best score on every metric: PESQ 3.88, STOI 0.96, CSIG 4.82, CBAK 3.85, COVL 4.82. It strictly dominates diffusion- and SB-based baselines (SGMSE+, SB-SE, SBCTM) at their larger step budgets, and exceeds the consistency-distilled ROSE-CD across all five quality metrics despite ROSE-CD's smaller K. The CBAK gain of +0.10 over the strongest SB baseline (SB-SE) and +0.48 over ROSE-CD indicates that dual-domain processing suppresses background noise more effectively than either pure-spectral or pure-waveform generative pipelines.

Ablation studies show performance drops following the asymmetric uncertainty design: removing the SB pathway causes the largest degradation (-0.63 PESQ), removing the MoE module yields-0.43 PESQ, and replacing uncertainty-aware fusion with equal weighting reduces performance by-0.17 PESQ. This ordering (0.63 > 0.43 > 0.17) supports the central design claim that ablations removing a distinct uncertainty channel are more damaging than those that only weaken fusion over intact pathways. Sequential variants in either direction also underperform parallel fusion (-0.30/-0.39 PESQ).

With only 8 sampling steps, HybridSB-MoE achieves an RTF of 0.28 (35 ms latency), giving a 4-5× speedup over SGMSE+ and SB-SE while surpassing their PESQ. The dual-domain design is most beneficial at low SNR (+0.13 over ROSE-CD at 0 dB). The calibration loss yields a fusion-weight ECE of 0.042, an order-of-magnitude reduction from the 0.12 achieved by an uncalibrated single-pathway baseline. Scene-stratified performance across all 14 noise types shows PESQ standard deviation < 0.03, confirming that combinatorial archetype coverage generalizes across scenes without per-noise tuning.

The paper concludes that the key to synergistic combination of SB and MoE lies in two linked ideas: pathway-typed asymmetric uncertainty fusion that selects between error regimes rather than averaging predictions, and a discretization bound that links the small-K inference budget to explicit training regularizers. Three counterfactual ablations show that no proper subset of components preserves the central design: the asymmetric fusion needs both a multi-expert path (for epistemic disagreement) and a stochastic bridge (for aleatoric variance), and the small-K guarantee needs both regularizers. Future work will extend the framework to multi-channel input, larger benchmarks (DNS, WHAMR!, CHiME), and finer NFE sweeps to empirically fit the predicted K(-α) rate.

Improvements for AI systems

Improvements to AI Systems:

  1. Adaptive Dual-Domain Processing: Implement a system that simultaneously processes inputs in two complementary domains (e.g., spectral and waveform for audio, or pixel and latent for images), with an asymmetric uncertainty fusion mechanism that dynamically weights each domain based on epistemic (model disagreement) and aleatoric (inherent noise) uncertainties, rather than fixed or equal weighting. This improves robustness across heterogeneous input conditions.

  2. Heterogeneous Mixture-of-Experts with Architectural Diversity: Replace homogeneous expert networks with a mixture of architecturally distinct experts (e.g., different receptive fields, kernel types, or inductive biases) routed via top-k selection. The epistemic uncertainty from expert disagreement then signals which inductive bias is failing, enabling targeted correction. This improves generalization to unseen noise/scene types without per-case tuning.

  3. Training-Informed Inference Budget Control: Introduce a theoretical discretization bound (analogous to Theorem 1) that links training regularizers (path-consistency and trajectory anchoring) to a guaranteed sampling error rate (e.g., K(-α) in Wasserstein distance). This allows the system to set a small, fixed number of inference steps (e.g., K=8) with an objective-level guarantee, rather than relying on empirical tuning, reducing latency and compute by 4-5x while maintaining or exceeding quality.

  4. Uncertainty-Calibrated Fusion: Add a calibration loss that anchors predicted epistemic and aleatoric uncertainties to actual reconstruction errors. This yields calibrated fusion weights (e.g., ECE reduced from 0.12 to 0.042), enabling the system to reliably switch between experts or domains based on trustworthy confidence signals—critical for safety-critical applications like medical imaging or autonomous driving.

  5. Cross-Timestep Consistency Regularization: Enforce agreement of intermediate predictions across different timesteps along the same generative trajectory (path-consistency) and anchor states to schedule-consistent reconstructions (trajectory loss). This improves stability of iterative refinement processes, reducing error accumulation and enabling fewer steps without quality loss.

What the Improved AI System Can Do:

  • Audio/Video Enhancement: Achieve state-of-the-art quality (e.g., PESQ 3.88, STOI 0.96) on noisy speech or video with only 8 sampling steps, running in real-time (RTF 0.28, 35 ms latency) on consumer hardware, while outperforming systems using 30-50 steps.

  • Robust Scene Adaptation: Handle diverse, unseen noise types (stationary hum, harmonic engine noise, non-stationary babble) with performance variance < 0.03 PESQ across 14 scene categories, without retraining or per-scene tuning.

  • Low-SNR Operation: Deliver +0.13 PESQ improvement over single-domain baselines at 0 dB SNR, making it suitable for extreme conditions like underwater acoustics or long-range communication.

  • Calibrated Decision-Making: Provide trustworthy confidence scores for each output, allowing downstream systems (e.g., hearing aids, voice assistants) to know when to rely on spectral vs. waveform processing, or when to escalate to human review.

  • Efficient Generative Modeling: Generate high-fidelity samples (images, audio, 3D shapes) with a fixed, small step budget, with a formal guarantee that error decreases predictably with more steps—enabling deployment on edge devices with strict latency constraints.

Sources

Related papers