Adapting Personalized Speech Enhancement for Low-Latency Audio-Visual Target-Speaker Extraction
eess.AS, cs.CV, cs.SD
Submitted: 2026-09-24
Updated: 2026-09-24
Code: https://github.com/resemble-ai/Resemblyzer
Terminology
Sources
- ICASSP 2023 Deep Noise Suppression Challenge
- LRS3-TED: a large-scale dataset for visual speech recognition
- Plug-and-Steer: Decoupling Separation and Selection in Audio-Visual Target Speaker Extraction
- Robust Speech Recognition via Large-Scale Weak Supervision
Related papers
- X-VC: Zero-shot Streaming Voice Conversion in Codec Space
- Autoregressive Guidance of Deep Spatially Selective Filters using Bayesian Tracking for Efficient Extraction of Moving Speakers
- Anonymization, Not Elimination: Utility-Preserved Speech Anonymization
- Towards Audio Token Compression in Large Audio Language Models
- WaveScat: Wavelet Scattering Front-Ends with Self-Supervised Features for Speech Deepfake Detection
- ProPS: Prompted Profile Synthesis for Natural Language-Conditioned Speaker Embedding Distributions