Language Orthogonalization for Zero-Shot Cross-Lingual Audio Deepfake Detection

arXiv:2609.16458 · eess.AS, cs.CL · Submitted 2026-09-15 · Read on arXiv

eess.AS, cs.CL

Submitted: 2026-09-15

Updated: 2026-09-15

Comments: Submitted to ICASSP 2027

License: http://creativecommons.org/licenses/by/4.0/

The gist: Audio deepfake detectors need to transfer to languages absent from training, as multilingual speech synthesis outpaces labeled anti-spoofing resources.

Terminology

Abstract

Audio deepfake detectors need to transfer to languages absent from training, as multilingual speech synthesis outpaces labeled anti-spoofing resources. While detectors increasingly rely on self-supervised speech models (S3Ms), these backbones encode language-dependent structure that confounds spoof cues. We address this confound through language orthogonalization, a target-free ridge map that removes S3M variation projected onto continuous language-identification (LID) embeddings. Across six languages, six S3M backbones, and all Leave-N-Out settings, it consistently reduces EER across unseen languages. Cross-lingual EER correlates with LID-space distance, where orthogonalization yields larger gains for more distant transfers.

Sources

Related papers