Beyond Short Segments: Expanding Speaker Embeddings with Vector Archives
eess.AS, cs.CL, cs.SD
Submitted: 2026-07-26
Updated: 2026-07-26
Comments: Accepted at INTERSPEECH 2026 (oral)
Code: https://github.com/slp-lab-research/vam_ecapa
License: http://creativecommons.org/licenses/by/4.0/
The gist: The performance of state-of-the-art speaker verification (SV) systems severely degrades on short utterances due to insufficient speaker-specific information.
Terminology
Abstract
The performance of state-of-the-art speaker verification (SV) systems severely degrades on short utterances due to insufficient speaker-specific information. To address this critical challenge, we propose the Vector Archive Mapping ECAPA (VAM-ECAPA), a novel system designed to enhance feature extraction from short-duration speech. The core of our system is the Transformer-based Vector Archive Mapping with Statistical Pooling (TVAMSP) module, which enriches information-scarce features by mapping them against a learnable Vector Archive of canonical speaker traits. By integrating the TVAMSP module into a strong WavLM+ECAPA-TDNN baseline, our system learns to map sparse features from short segments into robust, discriminative speaker representations. Experiments on the VoxCeleb1 benchmark show that our proposed VAM-ECAPA achieves a highly competitive EER of 8.334% on 1-second test segments, a 54.8% relative error reduction compared to a conventionally-trained baseline.
Sources
- SUPERB: Speech processing Universal PERformance Benchmark
- Exploring wav2vec 2.0 on speaker verification and language identification
- Segment Aggregation for short utterances speaker verification using raw waveforms
- Why does Self-Supervised Learning for Speech Recognition Benefit Speaker Recognition?
- Meta-Learning for Short Utterance Speaker Recognition with Imbalance Length Pairs
- Exploring the Encoding Layer and Loss Function in End-to-End Speaker and Language Recognition System
- Length- and Noise-aware Training Techniques for Short-utterance Speaker Recognition
- Neural Turing Machines
- ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification
- Attentive Statistics Pooling for Deep Speaker Embedding
- VoxCeleb2: Deep Speaker Recognition
- MUSAN: A Music, Speech, and Noise Corpus
- Efficient Implementation of the Room Simulator for Training Deep Neural Network Acoustic Models
Related papers
- X-VC: Zero-shot Streaming Voice Conversion in Codec Space
- Autoregressive Guidance of Deep Spatially Selective Filters using Bayesian Tracking for Efficient Extraction of Moving Speakers
- Anonymization, Not Elimination: Utility-Preserved Speech Anonymization
- Towards Audio Token Compression in Large Audio Language Models
- WaveScat: Wavelet Scattering Front-Ends with Self-Supervised Features for Speech Deepfake Detection
- ProPS: Prompted Profile Synthesis for Natural Language-Conditioned Speaker Embedding Distributions