Never Stop Speaking: a Denial-of-Service Attack on End-to-End Speech Language Models
Shuozhe Cheng, Kunlan Xiang, Mingxuan Li, Ji Zhang, Dongxiao Liu, Wenbo Jiang
University of Electronic Science and Technology of China
cs.SD, cs.AI
Submitted: 2026-08-11
Updated: 2026-08-12
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 75/100
The gist: This paper introduces a novel denial-of-service (DoS) attack targeting end-to-end (E2E) audio large language models (ALLMs).
Terminology
Summary
This paper introduces a novel denial-of-service (DoS) attack targeting end-to-end (E2E) audio large language models (ALLMs). The authors note that while DoS attacks are well-studied for text-based LLMs, they cannot be directly applied to E2E speech models because these models process continuous acoustic waveforms rather than discrete tokens. The paper states: "Existing text-based DoS attacks primarily rely on prompt engineering, such as adversarial suffixes or semantic inducement, which exploit the discrete nature of text inputs and therefore cannot be directly transferred to continuous speech inputs."
The proposed method is a white-box attack that optimizes imperceptible acoustic perturbations to force the model to generate excessively long outputs, thereby exhausting computational resources. The attack is formulated as a constrained optimization problem: min σ1,σ2,...,σn L(F(x̂))
subject to ∥σi∥p ≤ ϵi
for each voiced segment, where the adversarial example is constructed as x̂ = Concat(xi + σi i=1n, S).
The core of the method is a composite loss function with four components:
-
Weighted EOS logit loss (Leos): This suppresses the end-of-sentence token probability. Unlike simple summation, it assigns dynamic weights: "wt = whigh exp(t/N κ), if lEOS,t > 0; wt = 0, if lEOS,t ≤ 0." The gating component sets weight to zero for non-positive logits, and the exponential growth component emphasizes later autoregressive steps.
-
Top-k logit loss (Ltopk): This increases the probability of top-k tokens, indirectly suppressing EOS and preserving audio coherence.
-
Length loss (Llen): This encourages the expected generation length to approach the maximum limit, computed as
LLen = (Nmax − E[L])2 / Nmax.
-
Semantic alignment loss (Lsem): This maintains semantic consistency between original and perturbed audio by computing cosine similarity of speech encoder features.
To enhance stealthiness, the method uses Voice Activity Detection (VAD) to inject perturbations only into voiced regions of the audio, avoiding artifacts in silent regions.
Experiments were conducted on three open-source E2E ALLMs: LFM2.5-Audio-1.5B, Fun-Audio-Chat-8B, and Qwen2-Audio-7B-Instruct, using the OpenSLR and QCRI datasets. The attack uses Projected Gradient Descent (PGD) with 200 iterations per sample and an l∞ norm bound of ϵ = 10−4.
The main results show the attack achieves high success rates: On LFM2.5-Audio and FunAudioChat, our method achieves attack success rates of 87% and 84%, respectively, significantly outperforming all baseline methods.
The generated output length is extended to 941.88 and 920.24 tokens, which is more than four times longer than clean inputs.
The attack also increases GPU memory consumption: our method requiring 10.78 GB and 21.93 GB on LFM2.5-Audio and FunAudioChat, respectively, compared with only 8.89 GB and 17.26 GB for clean inputs.
Ablation studies reveal the importance of each loss component. Removing Leos causes a significant degradation in attack performance, reducing the ASR from 0.84 to 0.12 and the average output length from 950.24 to 289.43 tokens.
Removing Lsem slightly improves ASR and output length, but results in a substantial decline in response quality (from 4.75 to 3.92).
The VAD strategy is shown to maintain attack effectiveness while improving stealthiness: removing VAD results in a similar attack success rate (0.84) and slightly shorter output length... the response quality decreases from 4.75 to 4.52 without VAD.
The paper also evaluates robustness. The attack tolerates moderate additional noise, but lossy compression can affect the effectiveness of adversarial perturbations.
Under greedy decoding, the attack maintains strong performance with ASRs of 86%, 84%, and 85% on Liquid Audio, FunAudioChat, and Qwen2-Audio, respectively.
Cross-model transferability is limited, with transfer rates ranging from 7% to 13%.
The authors conclude that their method can effectively increase output lengths, achieve high attack success rates, and introduce substantial inference overhead while largely preserving the semantic information of the original speech,
highlighting the potential security risks of current E2E ALLMs and motivate the development of practical defense mechanisms against audio-domain adversarial DoS attacks.
Improvements for AI systems
Improvements to AI Systems:
-
Add adaptive output-length throttling with early-exit mechanisms. The improved system monitors autoregressive generation length in real time and triggers early stopping or switches to a more efficient decoding mode (e.g., speculative decoding) when output length exceeds a threshold relative to input length, preventing resource exhaustion from adversarial audio.
-
Implement input perturbation detection via voiced-region anomaly scoring. The system uses VAD to segment audio, then computes per-voiced-segment statistical deviations (e.g., spectral flatness, perturbation norm) from a learned clean-speech distribution. If anomalies exceed a threshold, the system flags the input as potentially adversarial and either rejects it or applies defensive denoising before inference.
-
Add EOS-logit robustness training. The system is fine-tuned with a regularizer that penalizes sharp drops in EOS probability during training, making the model less sensitive to small perturbations that suppress EOS. This increases resilience to the weighted EOS logit loss used in the attack.
-
Integrate cross-modal semantic consistency checks. The improved system computes cosine similarity between the speech encoder features of the input and a reconstructed/denoised version. If similarity drops below a threshold, the system reverts to a conservative decoding strategy (e.g., maximum length cap) or requests re-input, mitigating semantic-preserving attacks.
-
Deploy dynamic GPU memory management with output-length prediction. The system predicts expected output length from early decoding steps (using a lightweight regression head) and pre-allocates or limits memory accordingly, preventing the 25%+ memory overhead observed in the attack.
-
Add lossy-compression-based input sanitization. The system applies mild lossy compression (e.g., MP3 or Opus) to incoming audio before inference, which the paper shows degrades adversarial perturbations while preserving clean speech intelligibility, reducing attack success without harming normal users.
What the Improved AI System Can Do:
-
Resist adversarial DoS attacks by capping output length and detecting malicious perturbations in voiced regions, maintaining normal response times and GPU usage under attack.
-
Maintain high response quality for legitimate users, as semantic alignment and compression-based defenses do not degrade clean speech understanding.
-
Operate safely in real-time applications (e.g., voice assistants, transcription services) by preventing resource exhaustion and ensuring availability even when targeted by optimized acoustic perturbations.
-
Provide graceful degradation—when an attack is detected, the system can fall back to a text-only interface or a shorter-response mode, preserving core functionality.
Sources
- LFM2 Technical Report
- Qwen2-Audio Technical Report
- Ichigo: Mixed-Modal Early-Fusion Realtime Voice Assistant
- Denial-of-Service Poisoning Attacks against Large Language Models
- SlothSpeech: Denial-of-service Attack Against Speech Recognition Models
- Sparks of Large Audio Models: A Survey and Outlook
- ReasoningBomb: A Stealthy Denial-of-Service Attack by Inducing Pathologically Long Reasoning in Large Reasoning Models
- Excessive Reasoning Attack on Reasoning LLMs
- DrVoice: Parallel Speech-Text Voice Conversation Model via Dual-Resolution Speech Representations
- Crabs: Consuming Resource via Auto-generation for LLM-DoS Attack under Black-box Settings
- ExtendAttack: Attacking Servers of LRMs via Extending Reasoning
Related papers
- Few-Shot Open-Set Audio Classification via Transductive Prototype Refinement and Class Logit Enhancement
- Spectral Masking and Interpolation Attack (SMIA): A Black-box Adversarial Attack against Voice Authentication and Anti-Spoofing Systems
- AVMeme Exam: A Multimodal Multilingual Multicultural Benchmark for LLMs' Contextual and Cultural Knowledge and Thinking
- SoundWeaver: Compositional Warm-Starting for Text-to-Audio Diffusion Serving
- WASIL: In-the-Wild Arabic Spoken Interactions with LLMs
- Efficient Audiovisual Speech Processing via MUTUD: Multimodal Training and Unimodal Deployment