Speedbumps: Rejection Attacks on Speculative Decoding
summary
The gist
The gist Speculative Rejection Attacks (SRAs) are a novel class of attacks that cause draft and target models to disagree more often, resulting in fewer draft tokens being accepted per draft cycle,
In short
The research introduces Speedbump, an adversarial suffix attack that exploits how draft and target models interact during speculative decoding. By appending a malicious suffix, attackers can force the system to reject more draft tokens early in the process. This increases the number of required target model forward passes per generated token, effectively inflating inference costs without changing output quality.
Key concepts
- Speculative Decoding (SD)
- A technique that speeds up text generation by using a cheap 'draft' model to propose several candidate tokens at once. The main target is to verify these proposals quickly with a more powerful 'target' model in parallel. Speedbump attacks exploit the agreement between the draft and target models to make this verification process less efficient.
- Speedbump-P Proxy
- This proxy measures how well the target model supports the proposed draft tokens. It calculates acceptance probability based on which proposals are supported by the target model's support distribution. This helps define where in the sequence to place an adversarial suffix to maximize rejection.
- Speedbump-D Proxy
- This proxy focuses on measuring overlap between the full distributions of the draft and target models at each context. It calculates a measure of how much information is shared or aligned between what the draft model suggests and what the target model expects, guiding where to place an adversarial suffix for maximum disruption.
Terminology used across episodes
This episode discusses
- Speedbumps: Rejection Attacks on Speculative Decoding · Paper Radio
- Accelerating Large Language Model Decoding with Speculative Sampling
- OverThink: Slowdown Attacks on Reasoning LLMs · Paper Radio
- Inference with Reference: Lossless Acceleration of Large Language Models
- Remote Timing Attacks on Efficient Language Model Inference
- When Speculation Spills Secrets: Side Channels via Speculative Decoding In LLMs
- Excessive Reasoning Attack on Reasoning LLMs
- Mistletoe: Stealthy Acceleration-Collapse Attacks on Speculative Decoding
- Adversarial Prompts for Acceptance Collapse in Speculative Decoding · Paper Radio
- Universal and Transferable Adversarial Attacks on Aligned Language Models
- Training Verifiers to Solve Math Word Problems
- Detecting Language Model Attacks with Perplexity
The paper
Speedbumps: Rejection Attacks on Speculative Decoding · Read on arXiv
Adam Y. J. Jones, Yu Yuan, Sergio Maffeis
Imperial College London
Transcript
Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.
Nadia: Today's paper: "Speedbumps: Rejection Attacks on Speculative Decoding".
Elias: The gist Speculative Rejection Attacks (SRAs) are a novel class of attacks that cause draft and target models to disagree more often,
Nadia: First, who's behind it and why it matters.
Title and authors: Nadia: Let's talk about who did this. The paper is titled "Speedbumps: Rejection Attacks on Speculative Decoding," and it’s written by Adam Y. J. Jones, Yu Yuan, and Sergio Maffeis from Imperial College London.
Elias: They are researchers in AI security and cryptography, so they're looking at how these models work under attack or exploitation. What's important here is that the authors are focusing on how to actually exploit the mechanics of speculative decoding rather than just making generic noise.
Priya: As a privacy researcher, I’m curious if this means we need to worry about these attacks when we use speculative decoding for sensitive tasks, or if it’s mostly an issue for general speedup.
Nadia: It's not just general speedup anymore; the paper shows that even when people are using techniques like speculative decoding, there’s a vulnerability that can be exploited to increase the cost per token beyond what you would pay for standard autoregressive decoding without it.
Elias: They characterize these SRAs by saying they exploit the way draft and target models interact, specifically by causing them to disagree more frequently about the proposed tokens.
The paper's summary: Nadia: So, what exactly is the core mechanism of Speedbumps? It’s about appending a specific adversarial suffix to content controlled by an attacker, and this suffix is designed to degrade speculative decoding on a victim's prompt.
Elias: The attack optimizes the expected length of the accepted prefix by using two different proxies. One proxy, Speedbump-P, only looks at the drafted tokens themselves, while the other one, Speedbump-D, looks at both distributions—the draft and the target model’s distribution.
Priya: That sounds like they are trying to figure out if it’s better to mess with what the drafter proposes or what the target model expects overall. What do those proxies actually measure?
Nadia: Speedbump-P uses a Proposal Mass Proxy, which measures how much support the target model gives to each proposal presented for verification. Speedbump-D uses a Draft-Target Overlap Proxy, which looks at the minimum of the target probability and the overlap between the draft and target distributions.
Elias: The objective function they are optimizing is defined as JSpeedbump(x) = −one/Cx Σ c Xmc i Σ j=one αc j, which basically captures how reachable a prefix is by reducing acceptance at an early depth lowers the probability of reaching every subsequent depth.
The paper's improvements: Nadia: The authors show that Speedbump outperforms other approaches they compared it to, like Mistletoe and ADSD objectives on every single decoder they tested. They also found that weighting earlier output positions with a decay makes the Speedbump objective substantially stronger for certain models.
Elias: That suggests that focusing the attack on the beginning of the generated sequence is key to maximizing the slowdown effect, which makes sense because accepting tokens early affects all subsequent steps in generating a token.
Priya: If this is true, it means an attacker doesn't need to change a whole paragraph; they just need to mess with the very first few words of the prompt to start slowing down the process.
Nadia: Precisely. And they also looked at transferability, finding that Speedbump-P transfers strongly across different drafters on GSM8K, reducing performance by about one point two zero on average for other decoders compared to native suffixes which only reduce it by zero point seven three.
Elias: But the paper flags a limitation here too; Speedbump-D doesn't transfer well because it changes the drafter’s output, whereas Speedbump-P is more transferable across different drafters.
Conclusion: Nadia: So to wrap up, the main point of "Speedbumps: Rejection Attacks on Speculative Decoding" is that speculative decoding isn't safe under adversarial input because these SRAs can inflate per-token costs past what you would expect from a standard autoregressive process.
Elias: The authors show that this happens by appending an adversarial suffix to the input, and they give us two specific proxies, Speedbump-P and Speedbump-D, to understand how effective those attacks are.
Priya: From a measurement side, it’s clear that these attacks don't even need to alter the final response quality significantly for the attacker; they only need the cost increase. That means we have to be careful about relying on speculative savings when dealing with untrusted inputs in any application.
Nadia: Right. The paper concludes that LLM application operators shouldn't assume that speculative savings hold for untrusted input and should consider defenses against these Speculative Rejection Attacks.
Elias: They suggest countermeasures like detecting an attack by looking at the perplexity of the prompt, or by monitoring response perplexity or a drop in quality. And they also tested changing the target model itself as a way to defend against it.
Priya: It’s interesting that they found that switching the target model could actually reduce the effect of this cost inflation, which means model choice is part of the defense strategy.
Nadia: That’s what we're hearing about "Speedbumps: Rejection Attacks on Speculative Decoding." We've covered how these suffix attacks work, what they measure with those proxies, and why this matters for anyone using speculative decoding in a production environment.
More episodes
- 2610.10597-Certified Corruption Budgets: Anytime-Valid Leaderboard Claims under Adaptive Rigging
- 2610.10608-From Investigation Failures to Reliable SOC Agents: Understanding and Improving LLM-Based Alert Triage
- 2610.10612-PyCache Trap: The Inspection-Execution Gap in Agent Skill Scanners
- 2610.10644-SoK: Failure Modes in Common Criteria Product Evaluation - A Taxonomy and Design-for-Evaluability Guidance
- 2610.10617-MRCert: Towards Post-deployment Patch Robustness Certification for Adversarially Patched Samples via Type-specific Masking
- 2610.10620-When AI Finds Hidden Messages, Does It Report?
- 2610.10625-Safe at One Loop, Risky at Another: Aligning Safety Across Recurrent Depths in Looped Language Models
- 2610.10992-The Hint Weight of ML-DSA Signatures Is Key-Dependent: An Empirical Study across the Three FIPS 204 Parameter Sets
- 2610.10659-Applying Security by Design at the Point of Execution: How Governed Security Requirements Affect the Security of AI-Generated Code
- 2610.10735-DITTO: A Context-aware Pickle-based Pre-Trained Model Scanner for Effective Security Audits