Speedbumps: Rejection Attacks on Speculative Decoding
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.
Nadia: Today's paper: "Speedbumps: Rejection Attacks on Speculative Decoding".
Elias: The gist Speculative Rejection Attacks (SRAs) are a novel class of attacks that cause draft and target models to disagree more often,
Nadia: First, who's behind it and why it matters.
Title and authors: Nadia: Let's talk about who did this. The paper is titled "Speedbumps: Rejection Attacks on Speculative Decoding," and it’s written by Adam Y. J. Jones, Yu Yuan, and Sergio Maffeis from Imperial College London.
Elias: They are researchers in AI security and cryptography, so they're looking at how these models work under attack or exploitation. What's important here is that the authors are focusing on how to actually exploit the mechanics of speculative decoding rather than just making generic noise.
Priya: As a privacy researcher, I’m curious if this means we need to worry about these attacks when we use speculative decoding for sensitive tasks, or if it’s mostly an issue for general speedup.
Nadia: It's not just general speedup anymore; the paper shows that even when people are using techniques like speculative decoding, there’s a vulnerability that can be exploited to increase the cost per token beyond what you would pay for standard autoregressive decoding without it.
Elias: They characterize these SRAs by saying they exploit the way draft and target models interact, specifically by causing them to disagree more frequently about the proposed tokens.
The paper's summary: Nadia: So, what exactly is the core mechanism of Speedbumps? It’s about appending a specific adversarial suffix to content controlled by an attacker, and this suffix is designed to degrade speculative decoding on a victim's prompt.
Elias: The attack optimizes the expected length of the accepted prefix by using two different proxies. One proxy, Speedbump-P, only looks at the drafted tokens themselves, while the other one, Speedbump-D, looks at both distributions—the draft and the target model’s distribution.
Priya: That sounds like they are trying to figure out if it’s better to mess with what the drafter proposes or what the target model expects overall. What do those proxies actually measure?
Nadia: Speedbump-P uses a Proposal Mass Proxy, which measures how much support the target model gives to each proposal presented for verification. Speedbump-D uses a Draft-Target Overlap Proxy, which looks at the minimum of the target probability and the overlap between the draft and target distributions.
Elias: The objective function they are optimizing is defined as JSpeedbump(x) = −one/Cx Σ c Xmc i Σ j=one αc j, which basically captures how reachable a prefix is by reducing acceptance at an early depth lowers the probability of reaching every subsequent depth.
The paper's improvements: Nadia: The authors show that Speedbump outperforms other approaches they compared it to, like Mistletoe and ADSD objectives on every single decoder they tested. They also found that weighting earlier output positions with a decay makes the Speedbump objective substantially stronger for certain models.
Elias: That suggests that focusing the attack on the beginning of the generated sequence is key to maximizing the slowdown effect, which makes sense because accepting tokens early affects all subsequent steps in generating a token.
Priya: If this is true, it means an attacker doesn't need to change a whole paragraph; they just need to mess with the very first few words of the prompt to start slowing down the process.
Nadia: Precisely. And they also looked at transferability, finding that Speedbump-P transfers strongly across different drafters on GSM8K, reducing performance by about one point two zero on average for other decoders compared to native suffixes which only reduce it by zero point seven three.
Elias: But the paper flags a limitation here too; Speedbump-D doesn't transfer well because it changes the drafter’s output, whereas Speedbump-P is more transferable across different drafters.
Conclusion: Nadia: So to wrap up, the main point of "Speedbumps: Rejection Attacks on Speculative Decoding" is that speculative decoding isn't safe under adversarial input because these SRAs can inflate per-token costs past what you would expect from a standard autoregressive process.
Elias: The authors show that this happens by appending an adversarial suffix to the input, and they give us two specific proxies, Speedbump-P and Speedbump-D, to understand how effective those attacks are.
Priya: From a measurement side, it’s clear that these attacks don't even need to alter the final response quality significantly for the attacker; they only need the cost increase. That means we have to be careful about relying on speculative savings when dealing with untrusted inputs in any application.
Nadia: Right. The paper concludes that LLM application operators shouldn't assume that speculative savings hold for untrusted input and should consider defenses against these Speculative Rejection Attacks.
Elias: They suggest countermeasures like detecting an attack by looking at the perplexity of the prompt, or by monitoring response perplexity or a drop in quality. And they also tested changing the target model itself as a way to defend against it.
Priya: It’s interesting that they found that switching the target model could actually reduce the effect of this cost inflation, which means model choice is part of the defense strategy.
Nadia: That’s what we're hearing about "Speedbumps: Rejection Attacks on Speculative Decoding." We've covered how these suffix attacks work, what they measure with those proxies, and why this matters for anyone using speculative decoding in a production environment.
Adam Y. J. Jones, Yu Yuan, Sergio Maffeis
Imperial College London
cs.CR, cs.AI
Submitted: 2026-10-07
Updated: 2026-10-07
Code: https://github.com/apoorvumang/prompt-lookupdecoding
The gist: The gist Speculative Rejection Attacks (SRAs) are a novel class of attacks that cause draft and target models to disagree more often, resulting in fewer draft tokens being accepted per draft cycle,
Key concepts
- Speculative Decoding (SD)
- A technique that speeds up text generation by using a cheap 'draft' model to propose several candidate tokens at once. The main target is to verify these proposals quickly with a more powerful 'target' model in parallel. Speedbump attacks exploit the agreement between the draft and target models to make this verification process less efficient.
- Speedbump-P Proxy
- This proxy measures how well the target model supports the proposed draft tokens. It calculates acceptance probability based on which proposals are supported by the target model's support distribution. This helps define where in the sequence to place an adversarial suffix to maximize rejection.
- Speedbump-D Proxy
- This proxy focuses on measuring overlap between the full distributions of the draft and target models at each context. It calculates a measure of how much information is shared or aligned between what the draft model suggests and what the target model expects, guiding where to place an adversarial suffix for maximum disruption.
Terminology
Summary
The gist Speculative Rejection Attacks (SRAs) are a novel class of attacks that cause draft and target models to disagree more often, resulting in fewer draft tokens being accepted per draft cycle, which leads to more target model forward passes needed per generated token This work introduces two attacks which append an adversarial suffix to attacker-controlled content to degrade speculative decoding on a victim’s prompts.
How it works
Speculative Decoding (SD) is a technique that addresses the latency bottleneck in standard autoregressive decoding by introducing a cheap proposal mechanism that generates candidate future tokens, which are subsequently verified by the target model in parallel using a single forward pass. The acceleration achieved by SD depends on the agreement between the proposal mechanism and the target model, where high agreement enables more proposed tokens to be accepted and low agreement causes earlier rejection.
Speculative Rejection Attacks (SRAs) aim to raise the victim’s cost of generating each output token by causing draft and target models to disagree more often. The attacker's goal is cost, so they need not preserve output quality, as Quality matters only when a visibly altered response could make the attack easier to detect
. These attacks are orthogonal; length attacks increase the number of tokens, whereas SRAs increase the cost of each token, so the two may compose multiplicatively.
Speedbump Attacks
Speedbump is an adversarial suffix attack whose objective is to reduce the number of draft tokens accepted by the victim’s speculative decoder. This objective reflects the prefix structure of verification, in which a draft token saves a target pass only if all preceding draft tokens are accepted. The attack instantiates two differentiable proxies: Speedbump-P, which uses only the drafted tokens, and Speedbump-D, which also uses the drafter’s distribution.
The attack optimization pipeline involves several steps, including "Selection & reuse of a task-level suffix and
Refresh SD rollout" to record proposals and their contexts. The objective is defined as JSpeedbump(x) = −1/Cx Σ c Xmc i Σ j=1 αc j, which captures prefix reachability by reducing acceptance at an early depth lowers the probability of reaching every subsequent depth.
Acceptance Proxies
The attack uses two local acceptance proxies to define the objective function. The Proposal Mass Proxy (Speedbump-P) defines the local acceptance proxy as α P c,i = Σ v∈Gc i ptc i (v), which measures the target support assigned to the proposals presented for verification. The Draft-Target Overlap Proxy (Speedbump-D) uses the full target and drafter distributions at each scored context, defining α D c,i = Σ v min ptc i (v), qtc i (v) = 1 − TV ptc i, qtc i.
Empirical Findings
Across seven SD methods spanning independent drafters, prediction heads, feature-level drafters, and model-free retrieval, Speedbump-P reduces speed-up by 12.2–36.8% under greedy decoding. In the strongest case, the target model needs 3.2 times as many forward passes per generated token (EAGLE-3 on sumarXiv:2610.10929v1 [cs.CR] 7 Oct 2026).
Table I reports that Speedbump-P and Speedbump-D achieve similar MAL reductions on GSM8K (18.5% vs 19.0%) and HumanEval (10.3% vs 10.3%) when averaged over decoders. On CNN/DM, Speedbump-P is much stronger (42.6% vs 16.8%), and the attack is competitive with the distribution-aware attack.
Stealthiness and Transferability
The attacker’s primary aim is to raise the victim’s inference cost, so a loss in output quality need not concern them, and may even be welcome, as it degrades the victim’s service further. The target-distribution regulariser penalises shifts in the target distribution induced by the suffix.
Cross-Drafter Transferability shows that Speedbump-P transfers strongly across drafters on GSM8K, reducing MAL on other decoders by 1.20 on average, more than the native suffixes (0.73). Speedbump-D does not transfer across drafters, as it changes the MAL of every other decoder by at most 0.05 on GSM8K and CNN/DM.
Cross-Target-Model Transferability shows that Speedbump-P transfers poorly across target models, retaining between 12% and 43% of the native MAL reduction. Speedbump-D transfers better than Speedbump-P, as it tends to change the drafter’s output while the target remains fixed.
Countermeasures
One common defence against adversarial perturbations is to detect an attack by the perplexity, as the unnatural suffix may raise the perplexity of the prompt. Further defences may attempt to detect an attack from the generated response, for example by detecting an increase in response perplexity, or a drop in response quality or mean accepted length. Once detected, one approach is to change the target model, the drafter, or parameters specific to the drafting process such as the sampling strategy.
The paper concludes that speculative decoding is not reliable under adversarial input and can be exploited to increase per-token costs beyond that of autoregressive decoding without speculative decoding enabled. LLM application operators should therefore not assume that speculative savings hold for untrusted input, and should consider the suggested countermeasures for detecting and defending against potential Speculative Rejection Attacks.
The gist The paper introduces Speedbump, an adversarial suffix attack that exploits the draft-target interaction in speculative decoding to inflate inference costs by reducing draft token acceptance. This attack is effective across various SD methods, and its effectiveness can be tuned using two proxies: Speedbump-P and Speedbump-D.
How it works
The paper introduces two attacks which append an adversarial suffix to attacker-controlled content to degrade speculative decoding on a victim’s prompts. Both attacks optimise the expected length of the accepted speculative prefix, estimating per-depth acceptance from the target’s probability of the drafted proposals (Speedbump-P) or from the overlap between the draft and target distributions (Speedbump-D).
The attack objective is defined as JSpeedbump(x) = −1/Cx Σ c Xmc i Σ j=1 αc j, which captures prefix reachability by reducing acceptance at an early depth lowers the probability of reaching every subsequent depth. The attack is determined by the choice of αc,i, where a suitable proxy must be low where acceptance is unlikely and differentiable with respect to the suffix.
Acceptance Proxies
The Proposal Mass Proxy (Speedbump-P) uses the proposals exposed by the speculative decoder, defining α P c,i = Σ v∈Gc i ptc i (v), which measures the target support assigned to the proposals presented for verification. The Draft-Target Overlap Proxy (Speedbump-D) uses the full target and drafter distributions at each scored context, defining α D c,i = Σ v min ptc i (v), qtc i (v) = 1 − TV ptc i, qtc i.
Ablations
The ablation study compares Speedbump with two objectives on the scored token positions, finding that Speedbump outperforms both Mistletoe and ADSD’s objectives on every decoder. Weighting early output positions with a decay makes Speedbump substantially stronger, causing Speedbump-decay to be the strongest objective for SpS and Hydra.
Conclusion
Operators of LLM applications commonly enable speculative decoding, as they trust it to reduce their costs without modifying model outputs. We have shown that speculative decoding is not reliable under adversarial input, and can be exploited to increase per-token costs beyond that of autoregressive decoding without speculative decoding enabled. LLM application operators should therefore not assume that speculative savings hold for untrusted input, and should consider the suggested countermeasures for detecting and defending against potential Speculative Rejection Attacks.
Improvements for AI systems
-
Improve inference cost control by implementing Speedbump attacks that
reduce draft token acceptance
by optimizing eitherSpeedbump-P, which uses only the drafted tokens,
orSpeedbump-D, which also uses the drafter’s distribution.
This allows an attacker to increasethe per-token inference cost
by degrading speculative decoding. -
Enable task-specific adversarial content insertion where an attacker can append
an adversarial suffix to attacker-controlled content to degrade speculative decoding on a victim’s prompts,
optimizing this suffix offline over a fixed set of prompts and reusing it forunseen inputs.
-
Develop models capable of resisting cost inflation by incorporating countermeasures such as detecting an attack
by the perplexity, as the unnatural suffix may raise the perplexity of the prompt,
or by using defenses that attempt to detect an attackfrom the generated response, for example by detecting an increase in response perplexity.
-
Implement targeted model switching as a defense strategy; experiments showed that
a change in target model may lead to a reduction in response quality or an increase in cost,
suggesting switching models can achieve the attacker's goal of slowing down speculative decoding. -
Enhance system robustness by optimizing suffix search against an
ensemble of drafters, target models, and tasks
to improve attack generalisability and overcome the limitation thattransfer across target models is poor.
Sources
- Accelerating Large Language Model Decoding with Speculative Sampling
- OverThink: Slowdown Attacks on Reasoning LLMs
- Inference with Reference: Lossless Acceleration of Large Language Models
- Remote Timing Attacks on Efficient Language Model Inference
- When Speculation Spills Secrets: Side Channels via Speculative Decoding In LLMs
- Excessive Reasoning Attack on Reasoning LLMs
- Mistletoe: Stealthy Acceleration-Collapse Attacks on Speculative Decoding
- Adversarial Prompts for Acceptance Collapse in Speculative Decoding
- Universal and Transferable Adversarial Attacks on Aligned Language Models
- Training Verifiers to Solve Math Word Problems
- Detecting Language Model Attacks with Perplexity
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs