SMetric: Rethink LLM Scheduling for Serving Agents with Balanced Session-centric Scheduling
summary
The gist
This paper introduces SMetric, a novel LLM scheduling design specifically optimized for "agentic serving," where requests are issued by autonomous agents rather than humans.
In short
Researchers from Shanghai Jiao Tong University and Alibaba propose SMetric to optimize LLM scheduling for agentic workloads. Unlike human chats, AI agents rely heavily on KV cache reuse, often exceeding 80%. SMetric balances session-centric scheduling by prioritizing load balance for initial requests and memory reuse for subsequent ones, preventing cluster imbalance.
Key concepts
- Agentic serving
- Unlike traditional human-to-chatbot interactions, agentic serving involves programs or agents operating in high-speed machine loops. These agents require continuous blocks of text to make decisions, creating a constant and heavy workload that differs significantly from sporadic human questions.
- KV cache reuse
- This is the model's short-term memory of a conversation. While human chats have moderate reuse, AI agents frequently build on the same context, leading to token reuse rates exceeding 80%. Efficiently managing this memory is crucial for performance in agentic workflows.
- Balanced session-centric scheduling
- This method treats different parts of a conversation differently. It sends the first request of a new session to the least busy machine to ensure balance, then attempts to keep subsequent requests on the same machine to maximize memory reuse through KV cache affinity.
Terminology used across episodes
This episode discusses
- SMetric: Rethink LLM Scheduling for Serving Agents with Balanced Session-centric Scheduling · Paper Radio
- PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling
- SLOs-Serve: Optimized Serving of Multi-SLO LLMs
- iServe: An Intent-based Serving System for LLMs
- Harvest: Opportunistic Peer-to-Peer GPU Caching for LLM Inference
- Autellix: An Efficient Serving Engine for LLM Agents as General Programs
- Oneiros: KV Cache Optimization through Parameter Remapping for Multi-tenant LLM Serving
- DualPath: Breaking the Storage Bandwidth Bottleneck in Agentic LLM Inference
- DynaServe: Unified and Elastic Execution for Dynamic Disaggregated LLM Serving
- Stateful Large Language Model Serving with Pensieve
- DualMap: Enabling Both Cache Affinity and Load Balancing for Distributed LLM Serving
- DiffKV: Differentiated Memory Management for Large Language Models with Parallel KV Compaction
- SGLang: Efficient Execution of Structured Language Model Programs
- PolyServe: Efficient Multi-SLO Serving at Scale
The paper
SMetric: Rethink LLM Scheduling for Serving Agents with Balanced Session-centric Scheduling · Read on arXiv
Institute of Parallel and Distributed Systems, Shanghai Jiao Tong University · Alibaba Group · ShanghaiTech University
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "SMetric: Rethink LLM Scheduling for Serving Agents with Balanced Session-centric Scheduling".
Jane: The paper was written by Jiahao Wang, Kaizhan Lin, Kaixi Zhang, Jinbo Han, Chenguang Fang et al. from Institute of Parallel and Distributed Systems, Shanghai Jiao Tong University and Alibaba Group and ShanghaiTech University.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Jane, we have to talk about this new paper from the researchers at Shanghai Jiao Tong University and Alibaba. It’s called SM ETRIC: Rethink LLM Scheduling for Serving Agents with Balanced Session-centric Scheduling.
Jane: I was just looking at that title, Tom. It sounds like they are trying to solve a very specific problem that most people haven't even realized is happening yet.
Tom: Exactly, and they are focusing on "agentic serving," which is a huge shift from the way we usually think about LLMs.
Jane: Most of our current systems are built for humans typing into a chat box, but this paper says that's not the primary workload anymore.
Tom: Right, because an agent isn't a person waiting to read words; it's a program that needs a whole block of text to make its next move.
Jane: That changes everything about how we manage the hardware, doesn't it?
Lu: It really does, because we are moving from a world of sporadic human questions to a world of constant, high-speed machine loops.
Tom: Lu, do you think that changes the very way we design the clusters themselves?
Lu: I think it forces us to stop thinking about individual requests and start thinking about entire sessions of interaction.
Meng: That sounds like a massive headache for anyone trying to keep a cluster stable in production.
Jane: Why do you say that, Meng? Is it just the sheer volume?
Meng: It’s more than just volume; if these agents are looping so fast, any tiny mistake in how we route a request could cause a massive pile-up.
Lalam: I see it as a transition toward a more seamless digital ecosystem where the machines talk to each other without us even noticing.
Tom: So, instead of waiting for a human to click "send," these agents are essentially driving the entire traffic flow of the data center.
Jane: And if we don't rethink how we schedule that traffic, we’re going to end up with some very expensive, very idle GPUs.
Tom: We'll look at exactly what those findings were in our next segment.
Summary: Tom: We just touched on the shift in workload, but Jane, you can explain why this actually breaks current scheduling methods.
Jane: Well, it comes down to something called KV cache reuse, which is basically the model's short-term memory of the conversation.
Tom: And in a normal chat with a human, that memory isn't reused that much because humans change the subject or take long breaks.
Jane: But these agents are different; they keep building on the same context over and over again.
Tom: The paper actually found that in real-world agentic traces from BAILIAN, reuse exceeds eighty percent of the tokens.
Jane: That is a massive number compared to the fifty or sixty percent you see in regular chat.
Lu: That eighty percent figure is what caught my eye because it means the "memory" is almost always there if you know where to look.
Tom: But there's a catch, right? The paper mentions a major problem with how current schedulers handle all that reuse.
Lu: They are actually too obsessed with it; they try so hard to find the exact GPU that has the memory that they end up overloading just a few machines.
Meng: I can see why that would happen in my own deployment tests; if you always send a request to the same place, that machine is going to melt.
Jane: So, while they are trying to be efficient with memory, they are actually destroying the balance of the whole cluster?
Meng: Exactly, and you endre up with some GPUs running at a hundred percent while others are just sitting there doing nothing.
Lalam: It's a bit like a city where every taxi driver is trying to go to the exact same popular restaurant at once, leaving all the other streets empty.
Tom: That’s a perfect way to put it, Lalam.
Jane: So the researchers discovered that being "cache-aware" is actually causing a massive bottleneck.
Tom: We'll break down how they actually fixed this imbalance in the next part.
Improvements: Tom: So, we know the problem is that being too smart about memory makes the cluster imbalanced, but how does SM ETRIC fix it?
Jane: They use this clever idea called "balanced session-centric scheduling," which basically treats different parts of a conversation differently.
Tom: I love how they split it up; for the very first request in a new session, they don't care about memory at all and just focus on balance.
Jane: Right, they just send that first request to whatever machine is the least busy to make sure the workload is spread out from the start.
Tom: But then, for every request after that, they switch gears and try to keep it on the same machine to catch that eighty percent reuse.
Jane: It sounds like a balancing act, but they added these "guards" to make sure it doesn't go wrong.
Tom: They mentioned the `not overloaded` guard, which basically says "hey, if this machine is getting too crowded, forget the memory and just move to a new one."
Jane: And there’s also the `session not evicted` check, so they don't try to stick to a machine if the memory has already been cleared out.
Lu: I think that differential approach is brilliant because it acknowledges that the first step of a journey is different from the rest of the trip.
Meng: From an engineering standpoint, those guards are what make this actually viable for a real data center.
Tom: And the results are pretty staggering; they saw a ten to sixteen percent increase in total throughput under certain conditions.
Jane: They even saw prefill throughput jump by as much as thirty-four percent when they used a disaggregated architecture.
Meng: That’s a huge win for cost efficiency; if you can get thirty percent more work out of the same hardware, your margins change instantly.
Lu: It's not just about the money, though; it's about making these complex agentic reasoning loops possible in real-time.
Lalam: When we make these processes efficient, we enable a culture where digital assistants can actually handle deep work instead of just simple tasks.
Tom: We’re almost at the end here, so let’s wrap this all up.
Conclusion: Tom: We've covered a lot of ground with SM ETRIC: Rethink LLM Scheduling for Serving Agents with Balanced Session-centric Scheduling.
Jane: It really shows that as we move from humans to agents, our entire approach to managing AI clusters has to evolve.
Tom: We've seen how they moved from being too obsessed with memory to a much smarter, balanced way of handling sessions.
Lu: I’m honestly excited to see how this leads to even more autonomous systems that can run on much smaller, cheaper clusters.
Meng: My main takeaway is that this is a practical, deployable solution that actually addresses the reality of modern agent workloads.
Lalam: I think this is a step toward a world where AI isn't just a tool we prompt, but a reliable partner that can operate continuously in the background.
Tom: Thanks to everyone for joining us on the show today.
Jane: We'll see you next time with another fascinating paper!
More episodes
- 2610.10857-Self-Supervised Keyframe Discovery for Horizon-Invariant Behavior Cloning
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization