Spooftral: Can Voxtral Audio-Language Model Detect Speech Spoofing?
eess.AS, cs.CL
Submitted: 2026-09-23
Updated: 2026-09-23
Code: https://github.com/avishai111/Spooftral
Terminology
Sources
- A Survey of Threats Against Voice Authentication and Anti-Spoofing Systems
- Does Audio Deepfake Detection Generalize?
- Tandem spoofing-robust automatic speaker verification based on time-domain embeddings
- Improving Out-of-Domain Audio Deepfake Detection via Layer Selection and Fusion of SSL-Based Countermeasures
- Automatic speaker verification spoofing and deepfake detection using wav2vec 2.0 and data augmentation
- Audio-CoT: Exploring Chain-of-Thought Reasoning in Large Audio Language Model
- SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities
- SALMONN: Towards Generic Hearing Abilities for Large Language Models
- Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models
- Voxtral
- Audio-Language Models for Audio-Centric Tasks: A Systematic Survey
- Direct Language Model Alignment from Online AI Feedback
- Voxtral Realtime
- Ministral 3
- A Study of BFLOAT16 for Deep Learning Training
- StarGANv2-VC: A Diverse, Unsupervised, Non-parallel Framework for Natural-Sounding Voice Conversion
Related papers
- X-VC: Zero-shot Streaming Voice Conversion in Codec Space
- Autoregressive Guidance of Deep Spatially Selective Filters using Bayesian Tracking for Efficient Extraction of Moving Speakers
- Anonymization, Not Elimination: Utility-Preserved Speech Anonymization
- Towards Audio Token Compression in Large Audio Language Models
- WaveScat: Wavelet Scattering Front-Ends with Self-Supervised Features for Speech Deepfake Detection
- ProPS: Prompted Profile Synthesis for Natural Language-Conditioned Speaker Embedding Distributions