Generative Testing of Automated Speech Recognition Systems
Yanis Xabier Wilbrand Peña, Oliver Weißl, Andrea Stocco
cs.CR, cs.LG
Submitted: 2026-07-10
License: http://creativecommons.org/licenses/by/4.0/
The gist: Automatic speech recognition (ASR) systems have achieved high accuracy with transformer-based models, enabling deployment in critical applications.
Terminology
Abstract
Automatic speech recognition (ASR) systems have achieved high accuracy with transformer-based models, enabling deployment in critical applications. However, they remain vulnerable to adversarial manipulation, particularly in black-box settings where attacks must preserve perceptual naturalness. This work introduces GATAS, a black-box testing approach that generates failure inducing inputs by operating in the phoneme-level latent space of a text- to-speech model. Instead of perturbing waveforms directly, the approach interpolates latent representations to induce transcription errors while remaining within the manifold of natural speech. The attack is formulated as a multi-objective optimization problem balancing semantic divergence and perceptual quality. Our empirical evaluation against both white-box and black-box baselines shows that GATAS achieves a 98% success rate while producing lower distortion and higher perceptual quality, as confirmed by human studies. Despite operating without gradient access, GATAS remains competitive against white-box methods, highlighting that representation and perceptual alignment are more critical than access to model internals. Overall, our results demonstrate that untargeted latent-space optimization enables the efficient generation of realistic and effective test cases for ASR systems.
Sources
- Did you hear that? Adversarial Examples Against Automatic Speech Recognition
- Neural Predictor for Black-Box Adversarial Attacks on Speech Recognition
- Listen, Attend and Spell
- Feature-Aware Test Generation for Deep Learning Models
- Explaining and Harnessing Adversarial Examples
- Latent Regularization in Generative Test Input Generation
- There is more than one kind of robustness: Fooling Whisper with adversarial examples
- Muting Whisper: A Universal Acoustic Adversarial Attack on Speech Foundation Models
- Adversarial Attacks Against Automatic Speech Recognition Systems via Psychoacoustic Hiding
- Intriguing properties of neural networks
- Tacotron: Towards End-to-End Speech Synthesis
- HyperNet-Adaptation for Diffusion-Based Test Case Generation
- Generating Natural Adversarial Examples
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs