Language Orthogonalization for Zero-Shot Cross-Lingual Audio Deepfake Detection
eess.AS, cs.CL
Submitted: 2026-09-15
Updated: 2026-09-15
Comments: Submitted to ICASSP 2027
License: http://creativecommons.org/licenses/by/4.0/
The gist: Audio deepfake detectors need to transfer to languages absent from training, as multilingual speech synthesis outpaces labeled anti-spoofing resources.
Terminology
Abstract
Audio deepfake detectors need to transfer to languages absent from training, as multilingual speech synthesis outpaces labeled anti-spoofing resources. While detectors increasingly rely on self-supervised speech models (S3Ms), these backbones encode language-dependent structure that confounds spoof cues. We address this confound through language orthogonalization, a target-free ridge map that removes S3M variation projected onto continuous language-identification (LID) embeddings. Across six languages, six S3M backbones, and all Leave-N-Out settings, it consistently reduces EER across unseen languages. Cross-lingual EER correlates with LID-space distance, where orthogonalization yields larger gains for more distant transfers.
Sources
- ASVspoof 2019: Future Horizons in Spoofed and Fake Audio Detection
- ASVspoof 2021: Towards Spoofed and Deepfake Speech Detection in the Wild
- Are audio DeepFake detection models polyglots?
- When Spoof Detectors Travel: Evaluation Across 66 Languages in the Low-Resource Language Spoofing Corpus
- Self-Supervised Speech Models Encode Phonetic Context via Position-dependent Orthogonal Subspaces
- Scaling Self-Supervised Speech Models Uncovers Deep Linguistic Relationships: Evidence from the Pacific Cluster
- XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale
- mHuBERT-147: A Compact Multilingual HuBERT Model
- SEA-Spoof: Bridging The Gap in Multilingual Audio Deepfake Detection for South-East Asian
Related papers
- X-VC: Zero-shot Streaming Voice Conversion in Codec Space
- Autoregressive Guidance of Deep Spatially Selective Filters using Bayesian Tracking for Efficient Extraction of Moving Speakers
- Anonymization, Not Elimination: Utility-Preserved Speech Anonymization
- Towards Audio Token Compression in Large Audio Language Models
- WaveScat: Wavelet Scattering Front-Ends with Self-Supervised Features for Speech Deepfake Detection
- ProPS: Prompted Profile Synthesis for Natural Language-Conditioned Speaker Embedding Distributions