SMetric: Rethink LLM Scheduling for Serving Agents with Balanced Session-centric Scheduling
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "SMetric: Rethink LLM Scheduling for Serving Agents with Balanced Session-centric Scheduling".
Jane: The paper was written by Jiahao Wang, Kaizhan Lin, Kaixi Zhang, Jinbo Han, Chenguang Fang et al. from Institute of Parallel and Distributed Systems, Shanghai Jiao Tong University and Alibaba Group and ShanghaiTech University.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Jane, we have to talk about this new paper from the researchers at Shanghai Jiao Tong University and Alibaba. It’s called SM ETRIC: Rethink LLM Scheduling for Serving Agents with Balanced Session-centric Scheduling.
Jane: I was just looking at that title, Tom. It sounds like they are trying to solve a very specific problem that most people haven't even realized is happening yet.
Tom: Exactly, and they are focusing on "agentic serving," which is a huge shift from the way we usually think about LLMs.
Jane: Most of our current systems are built for humans typing into a chat box, but this paper says that's not the primary workload anymore.
Tom: Right, because an agent isn't a person waiting to read words; it's a program that needs a whole block of text to make its next move.
Jane: That changes everything about how we manage the hardware, doesn't it?
Lu: It really does, because we are moving from a world of sporadic human questions to a world of constant, high-speed machine loops.
Tom: Lu, do you think that changes the very way we design the clusters themselves?
Lu: I think it forces us to stop thinking about individual requests and start thinking about entire sessions of interaction.
Meng: That sounds like a massive headache for anyone trying to keep a cluster stable in production.
Jane: Why do you say that, Meng? Is it just the sheer volume?
Meng: It’s more than just volume; if these agents are looping so fast, any tiny mistake in how we route a request could cause a massive pile-up.
Lalam: I see it as a transition toward a more seamless digital ecosystem where the machines talk to each other without us even noticing.
Tom: So, instead of waiting for a human to click "send," these agents are essentially driving the entire traffic flow of the data center.
Jane: And if we don't rethink how we schedule that traffic, we’re going to end up with some very expensive, very idle GPUs.
Tom: We'll look at exactly what those findings were in our next segment.
Summary: Tom: We just touched on the shift in workload, but Jane, you can explain why this actually breaks current scheduling methods.
Jane: Well, it comes down to something called KV cache reuse, which is basically the model's short-term memory of the conversation.
Tom: And in a normal chat with a human, that memory isn't reused that much because humans change the subject or take long breaks.
Jane: But these agents are different; they keep building on the same context over and over again.
Tom: The paper actually found that in real-world agentic traces from BAILIAN, reuse exceeds eighty percent of the tokens.
Jane: That is a massive number compared to the fifty or sixty percent you see in regular chat.
Lu: That eighty percent figure is what caught my eye because it means the "memory" is almost always there if you know where to look.
Tom: But there's a catch, right? The paper mentions a major problem with how current schedulers handle all that reuse.
Lu: They are actually too obsessed with it; they try so hard to find the exact GPU that has the memory that they end up overloading just a few machines.
Meng: I can see why that would happen in my own deployment tests; if you always send a request to the same place, that machine is going to melt.
Jane: So, while they are trying to be efficient with memory, they are actually destroying the balance of the whole cluster?
Meng: Exactly, and you endre up with some GPUs running at a hundred percent while others are just sitting there doing nothing.
Lalam: It's a bit like a city where every taxi driver is trying to go to the exact same popular restaurant at once, leaving all the other streets empty.
Tom: That’s a perfect way to put it, Lalam.
Jane: So the researchers discovered that being "cache-aware" is actually causing a massive bottleneck.
Tom: We'll break down how they actually fixed this imbalance in the next part.
Improvements: Tom: So, we know the problem is that being too smart about memory makes the cluster imbalanced, but how does SM ETRIC fix it?
Jane: They use this clever idea called "balanced session-centric scheduling," which basically treats different parts of a conversation differently.
Tom: I love how they split it up; for the very first request in a new session, they don't care about memory at all and just focus on balance.
Jane: Right, they just send that first request to whatever machine is the least busy to make sure the workload is spread out from the start.
Tom: But then, for every request after that, they switch gears and try to keep it on the same machine to catch that eighty percent reuse.
Jane: It sounds like a balancing act, but they added these "guards" to make sure it doesn't go wrong.
Tom: They mentioned the `not overloaded` guard, which basically says "hey, if this machine is getting too crowded, forget the memory and just move to a new one."
Jane: And there’s also the `session not evicted` check, so they don't try to stick to a machine if the memory has already been cleared out.
Lu: I think that differential approach is brilliant because it acknowledges that the first step of a journey is different from the rest of the trip.
Meng: From an engineering standpoint, those guards are what make this actually viable for a real data center.
Tom: And the results are pretty staggering; they saw a ten to sixteen percent increase in total throughput under certain conditions.
Jane: They even saw prefill throughput jump by as much as thirty-four percent when they used a disaggregated architecture.
Meng: That’s a huge win for cost efficiency; if you can get thirty percent more work out of the same hardware, your margins change instantly.
Lu: It's not just about the money, though; it's about making these complex agentic reasoning loops possible in real-time.
Lalam: When we make these processes efficient, we enable a culture where digital assistants can actually handle deep work instead of just simple tasks.
Tom: We’re almost at the end here, so let’s wrap this all up.
Conclusion: Tom: We've covered a lot of ground with SM ETRIC: Rethink LLM Scheduling for Serving Agents with Balanced Session-centric Scheduling.
Jane: It really shows that as we move from humans to agents, our entire approach to managing AI clusters has to evolve.
Tom: We've seen how they moved from being too obsessed with memory to a much smarter, balanced way of handling sessions.
Lu: I’m honestly excited to see how this leads to even more autonomous systems that can run on much smaller, cheaper clusters.
Meng: My main takeaway is that this is a practical, deployable solution that actually addresses the reality of modern agent workloads.
Lalam: I think this is a step toward a world where AI isn't just a tool we prompt, but a reliable partner that can operate continuously in the background.
Tom: Thanks to everyone for joining us on the show today.
Jane: We'll see you next time with another fascinating paper!
Institute of Parallel and Distributed Systems, Shanghai Jiao Tong University · Alibaba Group · ShanghaiTech University
cs.DC, cs.AI
Submitted: 2026-07-09
Updated: 2026-09-13
Code: https://github.com/vllm-project/aibrix
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 89/100
The gist: This paper introduces SMetric, a novel LLM scheduling design specifically optimized for "agentic serving," where requests are issued by autonomous agents rather than humans.
Key concepts
- Agentic serving
- Unlike traditional human-to-chatbot interactions, agentic serving involves programs or agents operating in high-speed machine loops. These agents require continuous blocks of text to make decisions, creating a constant and heavy workload that differs significantly from sporadic human questions.
- KV cache reuse
- This is the model's short-term memory of a conversation. While human chats have moderate reuse, AI agents frequently build on the same context, leading to token reuse rates exceeding 80%. Efficiently managing this memory is crucial for performance in agentic workflows.
- Balanced session-centric scheduling
- This method treats different parts of a conversation differently. It sends the first request of a new session to the least busy machine to ensure balance, then attempts to keep subsequent requests on the same machine to maximize memory reuse through KV cache affinity.
Terminology
Summary
This paper introduces SMetric, a novel LLM scheduling design specifically optimized for agentic serving,
where requests are issued by autonomous agents rather than humans. It addresses the inefficiencies in existing schedulers that overly prioritize
KV reuse at the expense of load balance, a trade-off that ultimately caps cluster-wide throughput in agent-heavy workloads.
The unique characteristics of agentic workloads
The paper identifies that agentic serving
differs from human chat because agents act only on complete responses, making the cluster’s tokens per second (TPS) the primary goal.
Through a systematic study of real-world traces, the authors identify several key empirical observations:
-
KV reuse is dominant,
exceeding 80% in production traces. -
KV reuse follows session locality,
where intra-session reuse accounts for roughly 67% of reuses. -
First-turn requests enjoy KV reuse
via shared system prompts, accounting for 18–20% of reuses. -
KV is quickly reused,
with approximately 90% of reuses occurring within 100 seconds. -
Session token usage is skewed,
with the top 25% of sessions contributing over 80% of tokens.
Despite this skewness, the authors find that the tokens of sessions can still be balanced across serving instances over time
due to the high request-to-instance ratio in production.
How SMetric works
SMetric implements balanced session-centric scheduling
to resolve the tension between load balance and KV reuse. It employs a differential policy: for the first request of each session,
it routes purely for load balance to spread sessions across all instances; for follow-up requests,
it follows traditional cache-aware routing that prioritizes local-tier reuse. This approach is designed to be clean and stateless
by using session turn information as a routing hint, which can be derived from user inputs without the router needing to keep any record of users’ past requests.
To maintain efficiency and prevent imbalance, SMetric utilizes two specific guards:
-
not overloaded: This check prevents the scheduler from sticking to an instance that has become overloaded due to growing session lengths. -
session not evicted: This ensures that if a session's KV has been evicted from the local tier, the request is treated as a fresh session and balanced across instances.
Performance and evaluation
The authors evaluated SMetric on both PD-colocation and PD-disaggregation architectures using Qwen3 models. Under PD-colocation with a global store, SMetric improves cluster TPS by 10–16% over state-of-the-art schedulers. In the disaggregated setting, it improves prefill TPS by 2–34% across various global-tier provisionings.
Beyond throughput, SMetric also provides superior per-token latency:
-
It lowers the P50 and P90
time-per-output-token (TPOT)
under colocation. -
It achieves a lower
time-to-first-token (TTFT)
at all reported percentiles under disaggregation, including a 37% reduction at the median.
Improvements for AI systems
1. Differential Session-Centric Scheduler
-
Improvement: Implement a bifurcated routing policy that distinguishes between the first request of an agent session and all subsequent follow-up requests. The first request is routed using a pure load-balancing algorithm (e.g., least-loaded instance) to distribute sessions evenly across the cluster. Follow-up requests are routed using cache-aware logic to prioritize instances holding the session's previous KV cache.
-
Capability: The system maximizes cluster-wide Tokens Per Second (TPS) by preventing
hotspot
instances caused by shared system prompts, while simultaneously maintaining high efficiency through intra-session KV cache reuse.
2. Dynamic Load and Eviction Guardrails
-
Improvement: Integrate two pre-filter routing mechanisms into the scheduler: a
not overloadedguard that triggers a fallback to load-balancing if a target instance’s load exceeds a specific threshold of the cluster mean, and asession not evictedguard that reverts to load-balancing if the actual KV cache hit length is significantly lower than the expected history length. -
Capability: The system eliminates tail-latency spikes caused by long-running sessions accumulating on a single GPU and prevents the latency penalties associated with attempting to route requests to instances where the session's KV cache has already been evicted.
3. Stateless Metadata-Driven Routing
-
Improvement: Replace stateful session-to-instance mapping tables with a stateless inference mechanism that deduces the
session turn
number directly from the length of the conversation history provided in the standard LLM API request. -
Capability: The routing layer can scale to support massive numbers of concurrent agent sessions without the memory overhead, synchronization complexity, or garbage-collection requirements of maintaining a centralized session state database.
4. Optimized Prefill-Disaggregation Scheduling
-
Improvement: Apply the session-centric scheduling logic specifically to the prefill-tier of a disaggregated architecture (where prefill and decode tasks run on separate instance clusters).
-
Capability: The system maximizes prefill throughput (improving prefill TPS by up to 34%) by balancing the compute-intensive prefill workload across the cluster, which directly reduces Time-To-First-Token (TTFT) and prevents queueing bottlenecks in high-frequency agentic loops.
Sources
- PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling
- SLOs-Serve: Optimized Serving of Multi-SLO LLMs
- iServe: An Intent-based Serving System for LLMs
- Harvest: Opportunistic Peer-to-Peer GPU Caching for LLM Inference
- Autellix: An Efficient Serving Engine for LLM Agents as General Programs
- Oneiros: KV Cache Optimization through Parameter Remapping for Multi-tenant LLM Serving
- DualPath: Breaking the Storage Bandwidth Bottleneck in Agentic LLM Inference
- DynaServe: Unified and Elastic Execution for Dynamic Disaggregated LLM Serving
- Stateful Large Language Model Serving with Pensieve
- DualMap: Enabling Both Cache Affinity and Load Balancing for Distributed LLM Serving
- DiffKV: Differentiated Memory Management for Large Language Models with Parallel KV Compaction
- SGLang: Efficient Execution of Structured Language Model Programs
- PolyServe: Efficient Multi-SLO Serving at Scale
Related papers
- iScheduler: Reinforcement Learning-Driven Continual Optimization for Large-Scale Resource Investment Problems
- SAMM: Sharded Automated Market Maker
- InferScale: GPU-Native KV Injection for Personalized LLM Serving
- Vigil: Accountable Liveness against Selective Silence
- Steelhead: Interleaving Partially Synchronous and Asynchronous Commit Rules on a Shared DAG
- Pushing CPU Speech Synthesis to the Wall: Extreme Inference Tuning under Serverless Architecture and Billing