Beyond Short Segments: Expanding Speaker Embeddings with Vector Archives

arXiv:2609.25007 · eess.AS, cs.CL, cs.SD · Submitted 2026-07-26 · Read on arXiv

eess.AS, cs.CL, cs.SD

Submitted: 2026-07-26

Updated: 2026-07-26

Comments: Accepted at INTERSPEECH 2026 (oral)

Code: https://github.com/slp-lab-research/vam_ecapa

License: http://creativecommons.org/licenses/by/4.0/

The gist: The performance of state-of-the-art speaker verification (SV) systems severely degrades on short utterances due to insufficient speaker-specific information.

Terminology

Abstract

The performance of state-of-the-art speaker verification (SV) systems severely degrades on short utterances due to insufficient speaker-specific information. To address this critical challenge, we propose the Vector Archive Mapping ECAPA (VAM-ECAPA), a novel system designed to enhance feature extraction from short-duration speech. The core of our system is the Transformer-based Vector Archive Mapping with Statistical Pooling (TVAMSP) module, which enriches information-scarce features by mapping them against a learnable Vector Archive of canonical speaker traits. By integrating the TVAMSP module into a strong WavLM+ECAPA-TDNN baseline, our system learns to map sparse features from short segments into robust, discriminative speaker representations. Experiments on the VoxCeleb1 benchmark show that our proposed VAM-ECAPA achieves a highly competitive EER of 8.334% on 1-second test segments, a 54.8% relative error reduction compared to a conventionally-trained baseline.

Sources

Related papers