Deep Learning Based Relative Transfer Matrix Estimation for Multiple Sources and Multiple Microphones
Oshan A. B. Yalegama, Wageesha N. Manamperi
University of Moratuwa · The Australian National University
eess.AS, cs.AI, eess.SP
Submitted: 2026-08-12
Updated: 2026-08-13
Comments: Accepted to Interspeech 2026
Code: https://github.com/oshanyalegama/Denoised_
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 75/100
The gist: This paper investigates deep learning-based estimation of the Relative Transfer Matrix (ReTM), which was recently introduced as a generalization of the relative transfer function for multiple
Terminology
Summary
This paper investigates deep learning-based estimation of the Relative Transfer Matrix (ReTM), which was recently introduced as a generalization of the relative transfer function for multiple receivers and sources. The ReTM relates received signals between two sets of microphone groups with respect to all active sources in a room, and is independent of source signals but dependent on the spatial location of sources and the environment.
The authors propose three novel supervised learning frameworks for ReTM estimation:
-
STFT Convolutional Network (SCoNet): Applies depthwise convolution operations in the STFT domain. It inputs the STFT of microphone signals in group B and trains to estimate microphone signals in group A. SCoNet stacks real and imaginary parts along the channel dimension and performs two-dimensional depthwise convolution across time and channel axes, allowing the model to learn distinct parameter sets for each frequency component of the ReTM.
-
Convolutional Filter and Summation Network (FuSNet): Parameterizes the time-domain relationship using QA×QB learnable one-dimensional convolutional filters, whose weights correspond to the inverse Fourier transform of the ReTM elements. It inputs non-overlapping segments of length L with a context window of size 3L from time frames of microphone signals at group B.
-
LSTM-based Autoencoder Network (LAeNet): Uses a shared bidirectional-LSTM followed by a feedforward network. Each frequency bin of the concatenated STFTs is independently processed by a BiLSTM layer with shared weights, followed by layer normalization, time averaging, and a fully connected network to estimate ReTM coefficients.
The models are trained using a weighted sum of negative signal-to-distortion ratio (SDR) in the time domain and relative spectrum error (RSE) in the STFT domain, with weight factors α=1 and β=10.
Experiments were conducted in a 6×7×3 m rectangular room with T60=500 ms reverberation time, using Q=7 microphones (QA=3, QB=4) for scenarios A1 (two white Gaussian noise sources), A2 (two noise sources: air conditioner noise, music), and B (three sources: speech, air conditioner noise, music). Scenario C used Q=12 microphones (QA=5, QB=7) with the same sources as B.
Results show that FuSNet achieves the highest performance across most evaluation metrics in scenarios A1 and A2, particularly with uniform spectral sources (WGN). In scenarios B and C, all methods improve with more microphones, with SCoNet and FuSNet substantially outperforming the baseline covariance-based method. The baseline method is clearly outperformed by the proposed deep learning models, particularly in time domain measures like MSE and SDR.
In terms of computational efficiency, FuSNet demonstrates the lowest latency (1.07ms for QA=3, QB=4) and smallest parameter count (98.3k), while LAeNet exhibits the highest latency (1.80s) and largest parameter count (263.5k).
For speech enhancement application, LAeNet achieves the highest performance in both scenarios B and C, with an average SDR of +10.55 dB and average STOI improvement of 37%. SCoNet ranks second with average SDR of +8.41 dB and 33% STOI improvement. FuSNet's denoised speech retains significant echo noise, leading to lower performance, which the authors suggest could be improved through dereverberation algorithms.
The authors conclude that FuSNet consistently outperforms other approaches in ReTM estimation accuracy, while SCoNet and LAeNet achieve performance comparable to the baseline method. Future work aims to extend these models to source separation, speech dereverberation, and investigate optimal microphone grouping strategies.
Improvements for AI systems
Improvements to AI Systems:
-
Real-Time Multi-Source Spatial Filtering: Integrate FuSNet’s low-latency (1.07 ms) ReTM estimation into hearing aids or smart speakers. The improved system can dynamically separate and enhance multiple simultaneous sound sources (e.g., speech + music + noise) in reverberant rooms with minimal delay, enabling real-time augmented hearing.
-
Robust Speech Enhancement with Echo Suppression: Combine LAeNet’s high SDR (+10.55 dB) and STOI (+37%) gains with a post-dereverberation module (e.g., WPE) to remove residual echo noise. The improved system can deliver clean, intelligible speech in highly reverberant environments (T60=500 ms) for teleconferencing or voice assistants, outperforming current covariance-based methods.
-
Frequency-Aware Adaptive Beamforming: Use SCoNet’s per-frequency depthwise convolutions to learn frequency-dependent ReTM weights. The improved system can adaptively steer microphone arrays toward moving speakers while suppressing non-stationary noise (e.g., air conditioner hum) across different frequency bands, improving robustness in smart home devices.
-
Low-Compute Embedded Deployment: Replace the baseline covariance matrix inversion (high complexity) with FuSNet’s 98.3k-parameter, 1.07 ms model. The improved system can run on battery-powered IoT devices (e.g., smart glasses, wearables) to estimate spatial filters for source separation without cloud processing, preserving privacy and reducing energy consumption.
-
Multi-Microphone Generalization: Train SCoNet/FuSNet on variable microphone counts (e.g., QA=3, QB=4 to QA=5, QB=7) and use transfer learning. The improved system can automatically adapt to arbitrary microphone array geometries (e.g., phone, laptop, or distributed sensors) without re-training, enabling plug-and-play spatial audio processing.
-
Joint Source Separation and Dereverberation: Extend the LSTM-based LAeNet to output both ReTM and room impulse responses. The improved system can simultaneously separate overlapping speakers and remove late reverberation, improving automatic speech recognition accuracy in far-field scenarios (e.g., smart TVs, robot assistants).
Abstract
The Relative Transfer Matrix (ReTM), recently introduced as a generalization of the relative transfer function for multiple receivers and sources, shows promising performance when applied to speech enhancement in noisy environments. Estimating the ReTM of sound sources by exploiting the covariance matrices of multichannel recordings is highly beneficial for practical applications and, to date, remains the only proposed approach. This paper investigates deep learning-based ReTM estimation. We propose three novel supervised learning frameworks using time and short-time frequency transform domain convolutional networks, and a Long Short-Term Memory-based recurrent neural network. Experimental results demonstrate that the proposed models achieve more accurate estimation of the ReTM using five objective metrics compared to the covariance-based method. We also show the effectiveness of the proposed frameworks for speech enhancement, achieving performance on par with the baseline method.
Sources
- Relative Transfer Matrix Estimator using Covariance Subtraction
- Adam: A Method for Stochastic Optimization
Related papers
- X-VC: Zero-shot Streaming Voice Conversion in Codec Space
- Autoregressive Guidance of Deep Spatially Selective Filters using Bayesian Tracking for Efficient Extraction of Moving Speakers
- Anonymization, Not Elimination: Utility-Preserved Speech Anonymization
- Towards Audio Token Compression in Large Audio Language Models
- WaveScat: Wavelet Scattering Front-Ends with Self-Supervised Features for Speech Deepfake Detection
- ProPS: Prompted Profile Synthesis for Natural Language-Conditioned Speaker Embedding Distributions