Rethinking Latency Denial-of-Service: Attacking the LLM Serving Framework, Not the Model

arXiv:2602.07878 · cs.CR, cs.AI · Submitted 2026-02-08 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.

Nadia: Today's paper: "Rethinking Latency Denial-of-Service".

Elias: The gist: system-level optimization such as continuous batching provides a logical isolation to mitigate contagious latency impact on co-located users,

Nadia: First, who's behind it and why it matters.

Paper summary: Nadia: We started by looking at how latency attacks work in general. The authors point out that because LLM inference is so expensive, even small slowdowns can lead to big operating costs and availability risks for users.

Elias: They say that the existing research has mostly focused on algorithmic complexity attacks, crafting inputs to make the output length as bad as possible, but those are largely ineffective against current serving frameworks.

Priya: So they’re challenging the assumption that slow outputs from weird prompts are the main threat; they’re suggesting we look at what happens underneath the hood in how the service is running.

Nadia: They shift our focus to the system layer, introducing a new strategy called Fill and Squeeze. The thesis is that we should manipulate memory resources to trigger pathological behaviors in the scheduler rather than just trying to slow down one specific generation.

Elias: The core intuition behind this is manipulating memory itself to cause problems, not just using complex inputs to stall the math of the model. They break this down into two vectors: Fill and Squeeze.

Priya: Can you tell us more about what Fill actually does in that context? What is it trying to achieve with rapid resource exhaustion?

Nadia: The Fill vector focuses on rapidly exhausting the global KV cache usage. The authors suggest injecting adversarial requests that generate long sequences, sometimes using ambient traffic peaks, to quickly drive the cache usage to a point where it triggers the scheduler's admission control.

Elias: That admission control is what induces memory-based Head-of-Line blocking for subsequent users, essentially freezing their queue while others wait behind you.

Priya: And then we get to Squeeze, which exploits the preemption logic that’s already there in the system. What does forcing a loop between continuous preemption and recovery actually do to the system's resources?

Nadia: The Squeeze forces the scheduler into this oscillation, and it diverts precious GPU cycles into expensive operations like memory swapping or recomputation. It turns the resource management itself into a performance drain.

Elias: So, they’re essentially turning the way a server manages its queue and its memory as the actual vulnerability to be exploited by these new attacks in "Rethinking Latency Denial-of-Service: Attacking the LLM Serving Framework, Not the Model."

Priya: What this means for anyone who just listens to this show is that we’re moving from thinking about how hard it is to trick a model into slow computation to figuring out how to trick the infrastructure into mismanaging its own resources.

Conclusion: Nadia: We’ve talked about how this paper moves beyond just attacking the algorithm to targeting the serving framework, and we’re looking at how these authors frame that whole concept.

Elias: The title itself tells you a lot: they are specifically arguing that latency denial-of-service happens in the serving infrastructure, not because of flaws inside the model itself.

Priya: So, for someone listening who isn't deep in systems research, what does this really change about how we view security risks in AI deployment?

Nadia: It suggests that system-level optimizations like continuous batching aren't just nice features for efficiency; they are actually crucial defenses that provide logical isolation against contagious latency impacts between co-located users.

Elias: They show that these system mechanisms can actively mitigate the effects of latency attacks by isolating the impact on other users. It moves the responsibility from just hardening the model to hardening the environment running it.

Priya: The implication is that security researchers need to look at how these systems are designed to handle extreme load and memory pressure, not just what kind of inputs a model can process.

Nadia: Exactly. The Fill and Squeeze strategy demonstrates that by manipulating the scheduler’s state transitions via memory exhaustion and preemption loops, we can cause measurable performance degradation with a much lower attack cost than some existing methods.

Elias: It’s about showing that you can achieve higher latency on co-located users in a way that is significantly cheaper to execute compared to previous attacks, based on real-world pricing benchmarks.

Priya: So, the big picture here is that we need better understanding of how KV cache usage directly translates into user experience issues at the serving layer. We need to monitor those system metrics closely because they are what’s actually being exploited in this type of attack.

Zhejiang University

cs.CR, cs.AI

Submitted: 2026-02-08

Updated: 2026-10-08

Comments: NeurIPS 2026

Code: https://github.com/vllm-project/vllm

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 82/100

The gist: The gist: system-level optimization such as continuous batching provides a logical isolation to mitigate contagious latency impact on co-located users, revealing that existing model-centric latency

Key concepts

Continuous Batching
A system optimization where multiple user requests are processed together in a single batch. This technique helps isolate the latency impact of one slow request from others, preventing a single problematic user from blocking everyone else in the queue.
Fill Attack
This strategy involves injecting adversarial requests designed to rapidly exhaust available GPU memory by generating long sequences. The goal is to quickly push the global KV cache usage past a threshold that triggers scheduler admission control and causes Head-of-Line (HOL) blocking for other users.
Squeeze Attack
This attack exploits the scheduler's preemption logic by forcing it into a cycle of continuous preemption and recovery. This diverts valuable GPU cycles away from generation toward costly memory swapping or recomputation operations when the system approaches its saturation boundary.

Terminology

Summary

The gist: system-level optimization such as continuous batching provides a logical isolation to mitigate contagious latency impact on co-located users, revealing that existing model-centric latency attacks are largely ineffective against modern LLM serving systems.

How it works

The paper shifts the focus from the algorithm to the system layer by introducing a new Fill and Squeeze attack strategy targeting the deterministic state transitions of the scheduler The core intuition of our strategy is to manipulate memory resources to trigger pathological behaviors rather than simply slow down individual generations This attack operates via two vectors: 1) Fill, which focuses on rapid resource exhaustion by injecting adversarial requests that generate long sequences and potentially leveraging ambient traffic peaks By injecting adversarial requests that generate long sequences and potentially leveraging ambient traffic peaks, the attacker could quickly drive the global KV cache usage to a boundary point that triggers the scheduler’s admission control, inducing memory-based Head-of-Line (HOL) blocking that freezes the queue for subsequent users 2) Squeeze, which exploits the preemption logic by forcing the scheduler into an alternation between continuous preemption and recovery Once the system is brought to the saturation boundary, it forces the scheduler into an alternation between continuous preemption and recovery, that diverts precious GPU cycles into expensive memory swapping or recomputation operations

Side-Channel Probing and Cost Reduction

To implement this strategy in a production black-box environment, the paper develops a side-channel probing mechanism to estimate KV cache usage They discover that the Inter-Token Latency (ITL) of standard requests exhibits a strong correlation with global GPU memory usage due to the physical memory bandwidth contention They leverage this proxy to train a lightweight regressor such as lightGBM that allows the attacker to estimate memory boundaries without system privileges in black-box settings This ensures probing at a sub-millisecond scale, enabling near real-time synchronization with the attack plan

The Fill and Squeeze Strategy

The operational logic of the Fill and Squeeze attack is formalized in Algorithm 1 The algorithm dynamically toggles its injection strategy based on the calculated memory gap, ∆mem = Csat −Uˆsys, representing the remaining capacity before the scheduler forces victim eviction When the system is underutilized (∆mem > δlarge), the attacker enters the Fill regime, dispatching high-complexity prompts (p ∈ PHigh) to rapidly occupy VRAM and induce HOL blocking for subsequent user requests Conversely, as the memory state approaches the fragile preemption boundary (0 < ∆mem ≤ δsmall), the strategy shifts to the Squeeze regime, where minimal token injections (p ∈ PLow) are sufficient to tip the aggregate load over the physical threshold Csat, triggering a deterministic and computationally expensive preemption mechanism

Empirical Effectiveness and Cost-Efficiency

Extensive evaluations indicate by up to 20−280× average slowdown on Time to First Token and 1.5−4× average slowdown on Time Per Output Token compared to existing attacks with 30-40% lower attack cost The results show that our attack not only inflicts higher latency on co-located users with much lower cost based on real-world pricing benchmarks (e.g., GPT-4o) In a comparison table, F&S+plain-text achieved a 11.

Improvements for AI systems

  1. The AI system will implement a novel Fill and Squeeze attack strategy to exploit deterministic state transitions of the scheduler by manipulating memory resources, which is payload-agnostic that can be integrated with different latency attack methods. This allows the system to induce Head-of-Line blocking via the Fill phase and force a transition into repetitive preemption via the Squeeze phase.

  2. The system will utilize a side-channel probing mechanism, leveraging Inter-Token Latency (ITL) as a proxy for global GPU memory usage, allowing it to estimate KV-Cache usage in black-box settings by training a lightweight regressor like LightGBM. This enables the attack to be orchestrated with much less cost by identifying the system’s KV cache occupancy.

  3. The system will employ a constrained optimization formulation, minimizing the time-averaged cost of an attack while ensuring the condition: Exp. Memory ≥ Csat +δ, where Csat is the saturation point, thus forcing the scheduler into a state that causes system-wide performance degradation.

Sources

Related papers