Exploring Second-Order Pattern Recognition in Speaker Recognition
eess.AS, cs.AI
Submitted: 2026-09-10
Updated: 2026-09-25
Comments: Submit to ICASSP 2027
License: http://creativecommons.org/licenses/by/4.0/
The gist: In classical pattern recognition tasks, neural networks are trained to recognise human-defined patterns for model inputs.
Terminology
Abstract
In classical pattern recognition tasks, neural networks are trained to recognise human-defined patterns for model inputs. Some Explainable AI (XAI) methods can explain other latent patterns that underlie the network's recognition of inputs as human-defined patterns; in this work, we call these latent patterns second-order patterns, and we propose to discover them. To this end, we apply a hierarchical clustering algorithm to analyse whether representations learned by a speaker recognition network from utterances naturally form hierarchical clusters. Each resulting cluster represents a second-order pattern that characterises how the network recognises some known utterances as speaker identities. All the resulting second-order patterns are then semantically interpreted using the existing Hierarchical Cluster-Class Matching (HCCM) method. Furthermore, we propose a new task, second-order pattern recognition, to identify which discovered second-order patterns characterising known utterances are exhibited by an unseen utterance. To achieve this, we design the Hierarchical Cluster Navigation and Assignment (HCNA) method. HCNA recognises a known second-order pattern as applying to an unseen utterance when the unseen utterance's network representation lies within the extrapolation space of the cluster regarded as that second-order pattern. Our experiments show that the extrapolation mechanism introduced by HCNA substantially improves performance on the second-order pattern recognition task.
Sources
- Explainable AI in Speaker Recognition -- Making Latent Representations Understandable
- Explainable AI in Speaker Recognition -- Attention Map Visualisation and Evaluation
- Explainable Attribute-Based Speaker Verification
- Leveraging speaker attribute information using multi task learning for speaker verification and diarization
- Modern hierarchical, agglomerative clustering algorithms
Related papers
- X-VC: Zero-shot Streaming Voice Conversion in Codec Space
- Autoregressive Guidance of Deep Spatially Selective Filters using Bayesian Tracking for Efficient Extraction of Moving Speakers
- Anonymization, Not Elimination: Utility-Preserved Speech Anonymization
- Towards Audio Token Compression in Large Audio Language Models
- WaveScat: Wavelet Scattering Front-Ends with Self-Supervised Features for Speech Deepfake Detection
- ProPS: Prompted Profile Synthesis for Natural Language-Conditioned Speaker Embedding Distributions