Adversarial Prompts for Acceptance Collapse in Speculative Decoding
Run Wang, Chaoyi Zhou, Xi Liu, Yi Zhu, Amir Salarpour, Pedram MohajerAnsari, Zhi-Qi Cheng, Feng Luo, Siyu Huang, Mert D. Pesé
cs.CR, cs.CL, cs.LG
Submitted: 2026-07-23
License: http://creativecommons.org/licenses/by/4.0/
The gist: Lossless acceleration schemes, such as speculative decoding, promise significant inference speedups by relying on dynamic token-level alignment between a draft and a target model.
Terminology
Abstract
Lossless acceleration schemes, such as speculative decoding, promise significant inference speedups by relying on dynamic token-level alignment between a draft and a target model. However, this guarantee of semantic equivalence masks a severe operational vulnerability: draft-target alignment can be systematically attacked. In this paper, we introduce ADSD, which, to the best of our knowledge, is the first prompt-suffix attack that collapses verifier acceptance by pushing draft probability mass toward tokens the target is unlikely to accept. ADSD uses Soft-Collapse, a verifier-aligned surrogate derived from the asymmetric speculative acceptance rule, together with a target-preservation objective that discourages obvious task corruption. ADSD successfully generates highly effective adversarial suffixes. On the GSM8K dataset, our attack increases the mean sample time by 62.3% while preserving the task quality. We further show that this vulnerability exists across different domains, speculative decoding strategies, and model architectures.
Sources
- GPT-4 Technical Report
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Qwen Technical Report
- A Simple and Effective Pruning Approach for Large Language Models
- The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks
- Block Verification Accelerates Speculative Decoding
- LiteVLM: A Low-Latency Vision-Language Model Inference Pipeline for Resource-Constrained Environments
- Accelerating Large Language Model Decoding with Speculative Sampling
- SDSAT: Accelerating LLM Inference through Speculative Decoding with Semantic Adaptive Tokens
- Online Speculative Decoding
- Direct Alignment of Draft Model for Speculative Decoding with Chat-Fine-Tuned LLMs
- AdaEDL: Early Draft Stopping for Speculative Decoding of Large Language Models via an Entropy-based Lower Bound on Token Acceptance Probability
- Judge Decoding: Faster Speculative Sampling Requires Going Beyond Model Alignment
- When Speculation Spills Secrets: Side Channels via Speculative Decoding In LLMs
- LLM Safeguard is a Double-Edged Sword: Exploiting False Positives for Denial-of-Service Attacks
- Rethinking Latency Denial-of-Service: Attacking the LLM Serving Framework, Not the Model
- Amphista: Bi-directional Multi-head Decoding for Accelerating LLM Inference
- Speculative Decoding: Performance or Illusion?
- Universal and Transferable Adversarial Attacks on Aligned Language Models
- Detecting Language Model Attacks with Perplexity
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs