Rethinking Latency Denial-of-Service: Attacking the LLM Serving Framework, Not the Model

summary

Video file (mp4)

The gist

The gist: system-level optimization such as continuous batching provides a logical isolation to mitigate contagious latency impact on co-located users, revealing that existing model-centric latency

In short

The research shifts latency denial-of-service attacks from targeting individual LLM models to exploiting system-level scheduler behaviors. By manipulating memory resources through a 'Fill and Squeeze' strategy, attackers can induce pathological behaviors like Head-of-Line blocking or expensive preemption, effectively slowing down co-located users with significantly lower operational costs.

Key concepts

Continuous Batching
A system optimization where multiple user requests are processed together in a single batch. This technique helps isolate the latency impact of one slow request from others, preventing a single problematic user from blocking everyone else in the queue.
Fill Attack
This strategy involves injecting adversarial requests designed to rapidly exhaust available GPU memory by generating long sequences. The goal is to quickly push the global KV cache usage past a threshold that triggers scheduler admission control and causes Head-of-Line (HOL) blocking for other users.
Squeeze Attack
This attack exploits the scheduler's preemption logic by forcing it into a cycle of continuous preemption and recovery. This diverts valuable GPU cycles away from generation toward costly memory swapping or recomputation operations when the system approaches its saturation boundary.

Terminology used across episodes

This episode discusses

The paper

Rethinking Latency Denial-of-Service: Attacking the LLM Serving Framework, Not the Model · Read on arXiv

Zhejiang University

Transcript

Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.

Nadia: Today's paper: "Rethinking Latency Denial-of-Service".

Elias: The gist: system-level optimization such as continuous batching provides a logical isolation to mitigate contagious latency impact on co-located users,

Nadia: First, who's behind it and why it matters.

Paper summary: Nadia: We started by looking at how latency attacks work in general. The authors point out that because LLM inference is so expensive, even small slowdowns can lead to big operating costs and availability risks for users.

Elias: They say that the existing research has mostly focused on algorithmic complexity attacks, crafting inputs to make the output length as bad as possible, but those are largely ineffective against current serving frameworks.

Priya: So they’re challenging the assumption that slow outputs from weird prompts are the main threat; they’re suggesting we look at what happens underneath the hood in how the service is running.

Nadia: They shift our focus to the system layer, introducing a new strategy called Fill and Squeeze. The thesis is that we should manipulate memory resources to trigger pathological behaviors in the scheduler rather than just trying to slow down one specific generation.

Elias: The core intuition behind this is manipulating memory itself to cause problems, not just using complex inputs to stall the math of the model. They break this down into two vectors: Fill and Squeeze.

Priya: Can you tell us more about what Fill actually does in that context? What is it trying to achieve with rapid resource exhaustion?

Nadia: The Fill vector focuses on rapidly exhausting the global KV cache usage. The authors suggest injecting adversarial requests that generate long sequences, sometimes using ambient traffic peaks, to quickly drive the cache usage to a point where it triggers the scheduler's admission control.

Elias: That admission control is what induces memory-based Head-of-Line blocking for subsequent users, essentially freezing their queue while others wait behind you.

Priya: And then we get to Squeeze, which exploits the preemption logic that’s already there in the system. What does forcing a loop between continuous preemption and recovery actually do to the system's resources?

Nadia: The Squeeze forces the scheduler into this oscillation, and it diverts precious GPU cycles into expensive operations like memory swapping or recomputation. It turns the resource management itself into a performance drain.

Elias: So, they’re essentially turning the way a server manages its queue and its memory as the actual vulnerability to be exploited by these new attacks in "Rethinking Latency Denial-of-Service: Attacking the LLM Serving Framework, Not the Model."

Priya: What this means for anyone who just listens to this show is that we’re moving from thinking about how hard it is to trick a model into slow computation to figuring out how to trick the infrastructure into mismanaging its own resources.

Conclusion: Nadia: We’ve talked about how this paper moves beyond just attacking the algorithm to targeting the serving framework, and we’re looking at how these authors frame that whole concept.

Elias: The title itself tells you a lot: they are specifically arguing that latency denial-of-service happens in the serving infrastructure, not because of flaws inside the model itself.

Priya: So, for someone listening who isn't deep in systems research, what does this really change about how we view security risks in AI deployment?

Nadia: It suggests that system-level optimizations like continuous batching aren't just nice features for efficiency; they are actually crucial defenses that provide logical isolation against contagious latency impacts between co-located users.

Elias: They show that these system mechanisms can actively mitigate the effects of latency attacks by isolating the impact on other users. It moves the responsibility from just hardening the model to hardening the environment running it.

Priya: The implication is that security researchers need to look at how these systems are designed to handle extreme load and memory pressure, not just what kind of inputs a model can process.

Nadia: Exactly. The Fill and Squeeze strategy demonstrates that by manipulating the scheduler’s state transitions via memory exhaustion and preemption loops, we can cause measurable performance degradation with a much lower attack cost than some existing methods.

Elias: It’s about showing that you can achieve higher latency on co-located users in a way that is significantly cheaper to execute compared to previous attacks, based on real-world pricing benchmarks.

Priya: So, the big picture here is that we need better understanding of how KV cache usage directly translates into user experience issues at the serving layer. We need to monitor those system metrics closely because they are what’s actually being exploited in this type of attack.

More episodes

← Home