Explainable-by-Design Audio Deepfake Detection via Wiener-Hopf Linear Prediction
Mattia Tamiazzo, Simone Milani, Massimo Iuliani, Marco Fontani
cs.SD, cs.AI, cs.CR, cs.MM
Submitted: 2026-07-14
Comments: Accepted at ACM IH&MMSec 2026
License: http://creativecommons.org/licenses/by/4.0/
The gist: The rapid advancement of synthetic speech generation methods has made audio deepfake detection a critical challenge in multimedia forensics.
Terminology
Abstract
The rapid advancement of synthetic speech generation methods has made audio deepfake detection a critical challenge in multimedia forensics. While recent approaches achieve high detection accuracy, they typically rely on black-box architectures that offer limited interpretability and high computational complexity. In this paper, we propose an explainable-by-design audio deepfake detection framework based on Wiener-Hopf linear prediction, processed by a lightweight 2D Convolutional Neural Network (CNN). This design enables a direct and transparent connection between classification outcomes and the acoustic properties of the signal. Experimental results on benchmark datasets demonstrate competitive detection performance while maintaining significantly lower computational complexity compared to state-of-the-art solutions. The interpretability analysis using Grad-CAM reveals that the classifier focuses on low-order predictor coefficients and on silence and transitional regions, suggesting that the Wiener-Hopf predictor captures reverberation characteristics and subtle statistical inconsistencies in synthetic speech. Finally, robustness experiments show that fine-tuning effectively recovers detection performance under common post-processing degradations, including additive noise, MP3 compression, and telephone filtering.
Sources
- Multi-Task Transformer for Explainable Speech Deepfake Detection via Formant Modeling
- Automatic speaker verification spoofing and deepfake detection using wav2vec 2.0 and data augmentation
- ASVspoof 2019: Future Horizons in Spoofed and Fake Audio Detection
- RawTFNet: A Lightweight CNN Architecture for Speech Anti-spoofing
- Audio Deepfake Detection: A Survey
Related papers
- Few-Shot Open-Set Audio Classification via Transductive Prototype Refinement and Class Logit Enhancement
- Spectral Masking and Interpolation Attack (SMIA): A Black-box Adversarial Attack against Voice Authentication and Anti-Spoofing Systems
- AVMeme Exam: A Multimodal Multilingual Multicultural Benchmark for LLMs' Contextual and Cultural Knowledge and Thinking
- SoundWeaver: Compositional Warm-Starting for Text-to-Audio Diffusion Serving
- WASIL: In-the-Wild Arabic Spoken Interactions with LLMs
- Efficient Audiovisual Speech Processing via MUTUD: Multimodal Training and Unimodal Deployment