Rethinking Latency Denial-of-Service: Attacking the LLM Serving Framework, Not the Model
summary
The gist
The gist: system-level optimization such as continuous batching provides a logical isolation to mitigate contagious latency impact on co-located users, revealing that existing model-centric latency
In short
The research shifts latency denial-of-service attacks from targeting individual LLM models to exploiting system-level scheduler behaviors. By manipulating memory resources through a 'Fill and Squeeze' strategy, attackers can induce pathological behaviors like Head-of-Line blocking or expensive preemption, effectively slowing down co-located users with significantly lower operational costs.
Key concepts
- Continuous Batching
- A system optimization where multiple user requests are processed together in a single batch. This technique helps isolate the latency impact of one slow request from others, preventing a single problematic user from blocking everyone else in the queue.
- Fill Attack
- This strategy involves injecting adversarial requests designed to rapidly exhaust available GPU memory by generating long sequences. The goal is to quickly push the global KV cache usage past a threshold that triggers scheduler admission control and causes Head-of-Line (HOL) blocking for other users.
- Squeeze Attack
- This attack exploits the scheduler's preemption logic by forcing it into a cycle of continuous preemption and recovery. This diverts valuable GPU cycles away from generation toward costly memory swapping or recomputation operations when the system approaches its saturation boundary.
Terminology used across episodes
This episode discusses
- Rethinking Latency Denial-of-Service: Attacking the LLM Serving Framework, Not the Model · Paper Radio
- Detecting Language Model Attacks with Perplexity
- Evaluating Large Language Models Trained on Code
- An Engorgio Prompt Makes Large Language Model Babble on
- BenchOverflow: Measuring Overflow in Large Language Models via Plain-Text Prompts
- LingoLoop Attack: Trapping MLLMs via Linguistic Context and State Entrapment into Endless Loops
- Ascendra: Dynamic Request Prioritization for Efficient LLM Serving
- Baseline Defenses for Adversarial Attacks Against Aligned Language Models
- POT: Inducing Overthinking in LLMs via Black-Box Iterative Optimization
- LLM Self Defense: By Self Examination, LLMs Know They Are Being Tricked
- DynaServe: Unified and Elastic Execution for Dynamic Disaggregated LLM Serving
- Excessive Reasoning Attack on Reasoning LLMs
- Gemma: Open Models Based on Gemini Research and Technology
- LLaMA: Open and Efficient Foundation Language Models
- Fast Distributed Inference Serving for Large Language Models
- Pie: Pooling CPU Memory for LLM Inference
- BitHydra: Towards Bit-flip Inference Cost Attack against Large Language Models
- Qwen3 Technical Report
- LLM Inference Unveiled: Survey and Roofline Model Insights
- Time Will Tell: Timing Side Channels via Output Token Count in Large Language Models
- PD 3F: A Pluggable and Dynamic DoS-Defense Framework Against Resource Consumption Attacks Targeting Large Language Models
The paper
Rethinking Latency Denial-of-Service: Attacking the LLM Serving Framework, Not the Model · Read on arXiv
Zhejiang University
Transcript
Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.
Nadia: Today's paper: "Rethinking Latency Denial-of-Service".
Elias: The gist: system-level optimization such as continuous batching provides a logical isolation to mitigate contagious latency impact on co-located users,
Nadia: First, who's behind it and why it matters.
Paper summary: Nadia: We started by looking at how latency attacks work in general. The authors point out that because LLM inference is so expensive, even small slowdowns can lead to big operating costs and availability risks for users.
Elias: They say that the existing research has mostly focused on algorithmic complexity attacks, crafting inputs to make the output length as bad as possible, but those are largely ineffective against current serving frameworks.
Priya: So they’re challenging the assumption that slow outputs from weird prompts are the main threat; they’re suggesting we look at what happens underneath the hood in how the service is running.
Nadia: They shift our focus to the system layer, introducing a new strategy called Fill and Squeeze. The thesis is that we should manipulate memory resources to trigger pathological behaviors in the scheduler rather than just trying to slow down one specific generation.
Elias: The core intuition behind this is manipulating memory itself to cause problems, not just using complex inputs to stall the math of the model. They break this down into two vectors: Fill and Squeeze.
Priya: Can you tell us more about what Fill actually does in that context? What is it trying to achieve with rapid resource exhaustion?
Nadia: The Fill vector focuses on rapidly exhausting the global KV cache usage. The authors suggest injecting adversarial requests that generate long sequences, sometimes using ambient traffic peaks, to quickly drive the cache usage to a point where it triggers the scheduler's admission control.
Elias: That admission control is what induces memory-based Head-of-Line blocking for subsequent users, essentially freezing their queue while others wait behind you.
Priya: And then we get to Squeeze, which exploits the preemption logic that’s already there in the system. What does forcing a loop between continuous preemption and recovery actually do to the system's resources?
Nadia: The Squeeze forces the scheduler into this oscillation, and it diverts precious GPU cycles into expensive operations like memory swapping or recomputation. It turns the resource management itself into a performance drain.
Elias: So, they’re essentially turning the way a server manages its queue and its memory as the actual vulnerability to be exploited by these new attacks in "Rethinking Latency Denial-of-Service: Attacking the LLM Serving Framework, Not the Model."
Priya: What this means for anyone who just listens to this show is that we’re moving from thinking about how hard it is to trick a model into slow computation to figuring out how to trick the infrastructure into mismanaging its own resources.
Conclusion: Nadia: We’ve talked about how this paper moves beyond just attacking the algorithm to targeting the serving framework, and we’re looking at how these authors frame that whole concept.
Elias: The title itself tells you a lot: they are specifically arguing that latency denial-of-service happens in the serving infrastructure, not because of flaws inside the model itself.
Priya: So, for someone listening who isn't deep in systems research, what does this really change about how we view security risks in AI deployment?
Nadia: It suggests that system-level optimizations like continuous batching aren't just nice features for efficiency; they are actually crucial defenses that provide logical isolation against contagious latency impacts between co-located users.
Elias: They show that these system mechanisms can actively mitigate the effects of latency attacks by isolating the impact on other users. It moves the responsibility from just hardening the model to hardening the environment running it.
Priya: The implication is that security researchers need to look at how these systems are designed to handle extreme load and memory pressure, not just what kind of inputs a model can process.
Nadia: Exactly. The Fill and Squeeze strategy demonstrates that by manipulating the scheduler’s state transitions via memory exhaustion and preemption loops, we can cause measurable performance degradation with a much lower attack cost than some existing methods.
Elias: It’s about showing that you can achieve higher latency on co-located users in a way that is significantly cheaper to execute compared to previous attacks, based on real-world pricing benchmarks.
Priya: So, the big picture here is that we need better understanding of how KV cache usage directly translates into user experience issues at the serving layer. We need to monitor those system metrics closely because they are what’s actually being exploited in this type of attack.
More episodes
- 2610.10644-SoK: Failure Modes in Common Criteria Product Evaluation - A Taxonomy and Design-for-Evaluability Guidance
- 2610.10617-MRCert: Towards Post-deployment Patch Robustness Certification for Adversarially Patched Samples via Type-specific Masking
- 2610.10620-When AI Finds Hidden Messages, Does It Report?
- 2610.10625-Safe at One Loop, Risky at Another: Aligning Safety Across Recurrent Depths in Looped Language Models
- 2610.10992-The Hint Weight of ML-DSA Signatures Is Key-Dependent: An Empirical Study across the Three FIPS 204 Parameter Sets
- 2610.10659-Applying Security by Design at the Point of Execution: How Governed Security Requirements Affect the Security of AI-Generated Code
- 2610.10735-DITTO: A Context-aware Pickle-based Pre-Trained Model Scanner for Effective Security Audits
- 2610.10742-BRANCH: Bypassing Multi-Scanner AI Guardrails
- 2610.10752-Detection-Guided Adaptive Purification with Diffusion Models for Robust Audio Deepfake Detection
- 2610.10766-CPU-Auth: Device Fingerprinting for Authentication via DVFS Side-Channel