Astrolabe: Balancing Load in LLM Serving with Randomized Prediction-Guided Scheduling

summary

Video file (mp4)

The gist

This paper presents Astrolabe, a randomized prediction-guided scheduler for one-shot request dispatch in multi-instance large language model (LLM) serving.

This episode discusses

The paper

Astrolabe: Balancing Load in LLM Serving with Randomized Prediction-Guided Scheduling · Read on arXiv

Wei Da, Evangelia Kalyvianaki

University of Cambridge

This paper presents Astrolabe, a randomized prediction-guided scheduler for one-shot request dispatch in multi-instance large language model (LLM) serving. Astrolabe improves load balancing without relying on migration-based rebalancing, whose KV-cache transfers can introduce substantial overhead and network contention under high load. It combines response-length estimation, per-instance simulation-based latency prediction, and a power-of-two-choices dispatch policy to balance load while avoiding request herding. On the default Llama-2-7B/ShareGPT setup, Astrolabe matches the SLO capacity of the best load-aware baseline (31.6 versus 31.5 QPS), while reducing mean time-to-first-token (TTFT) by 8 to 36 percent, P99 TTFT by 16 to 77 percent, and mean end-to-end (E2E) latency by up to 5.6 percent, with approximately six times fewer preemptions once capacity is reached. Under configuration shifts, Astrolabe improves SLO capacity by up to 6 percent on Qwen2-7B and 7.1 percent under tight batching, reduces mean E2E latency by 6 to 9 percent under bursty arrivals, and achieves an approximately 2.8-fold reduction in per-predictor CPU usage relative to full fanout. With migration enabled on A100 GPUs, Astrolabe outperforms Llumnix by up to 2.6 times in throughput while achieving orders-of-magnitude lower per-token latency at saturation.

DOI: 10.1145/3793230.3837768

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Astrolabe: Balancing Load in LLM Serving with Randomized Prediction-Guided Scheduling".

Jane: The paper was written by Wei Da and Evangelia Kalyvianaki from University of Cambridge.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title and Authors: Tom: Welcome back to the show, everyone! I'm Tom, and joining me as always is the brilliant Jane. Today we're cracking open a fresh arXiv paper that's got the systems community buzzing. It's called "Astrolabe: Balancing Load in LLM Serving with Randomized Prediction-Guided Scheduling."

Jane: And I'm Jane! Tom, I have to say, I love this title. An astrolabe was an ancient device used by sailors to navigate by the stars, right? It helped you figure out where you were and where to go. And this paper is doing exactly that for large language model servers—helping them navigate where to send requests.

Tom: That's a perfect way to put it, Jane. The authors are Wei Da and Evangelia Kalyvianaki from the University of Cambridge. And they're tackling a problem that anyone running a big AI service feels every single day. When you have a cluster of GPUs serving chatbots or code generators, how do you decide which GPU gets which user request?

Jane: Right, and the obvious answer seems easy—just spread them out evenly, right? But it's not that simple. The paper explains that LLM inference is just wildly unpredictable. You don't know how long a response will be when a request comes in. One user asks for a haiku, another asks for a full essay. And that uncertainty wreaks havoc on load balancing.

Tom: Exactly. The old-school approach is like round-robin, just sending requests to servers in a circle. But that ignores the fact that some requests are heavy and some are light. It's like a grocery store where every cashier gets the same number of customers, but one customer has a cart overflowing while another has a single loaf of bread. The lines get uneven fast.

Jane: And the more sophisticated systems try to fix this by migrating requests between servers mid-flight. That means moving the "KV cache"—the memory of what the model has already computed—across the network. But that transfer is expensive and can clog up the network, especially when things get busy.

Tom: So Astrolabe's big idea is to do the smart thing *before* sending the request, not after. They predict how long the response will be, they simulate what the latency would be on a couple of candidate servers, and then they send the request to the best one. No migration needed.

Jane: It's proactive instead of reactive. And the "randomized" part in the title is clever too. Instead of checking every server, they just check two random ones and pick the better. It's a classic trick from distributed systems called "power of two choices," and it prevents a stampede where every request tries to go to the same "best" server.

Tom: And the results are pretty stunning. On their test cluster, they matched the capacity of the best existing system but cut the time-to-first-token by up to thirty-six percent and reduced preemptions by six times. We'll get into the nitty-gritty of those numbers in a bit.

Jane: So this isn't just a theoretical idea. They built it, they tested it against real baselines, and it wins. I'm excited to dig into how the prediction actually works.

Tom: Stick around—that's exactly what we're covering next.

Summary: Tom: Welcome back. We're diving into "Astrolabe: Balancing Load in LLM Serving with Randomized Prediction-Guided Scheduling." Last time we set the scene—the problem of unpredictable request lengths and the cost of migration. Now let's talk about the actual machinery. Jane, how does Astrolabe pull off this prediction trick?

Jane: So there are three main pieces working together, Tom. First, there's a "length tagger." It's a small, fast machine-learning model—a fine-tuned RoBERTa—that looks at the user's prompt and guesses how many tokens the response will be. It's not perfect, but it's pretty good, with a mean absolute error of about seventy-nine tokens.

Tom: And that guess feeds into the second piece, the "Predictor" sidecar that runs on each server. This is the really cool part. Each server has a little simulator that knows its own current state—how many requests are running, how much memory is free, what's in the queue. When a new request comes in, the simulator plays it forward against that real state and estimates the end-to-end latency.

Jane: Right, and they built this on top of Vidur, which is an existing LLM serving simulator. But they had to rework it to run online, in real time, instead of offline for planning. They made it fast enough to give a prediction in milliseconds.

Tom: And the third piece is the global scheduler, the "brain" that makes the final call. When a request arrives, the scheduler picks two random servers, asks their Predictors for a latency estimate, and sends the request to whichever one predicts a lower latency. That's the power-of-two choices strategy.

Jane: And it's worth emphasizing why that randomness matters. If the scheduler always picked the absolute best server according to the predictions, then under a burst of traffic, every request would see the same "best" answer and pile onto one server. That's called herding. By sampling randomly, they break that pattern.

Tom: The paper compares this against five other schedulers, including round-robin and the heuristic used by the Llumnix system. And across the board, Astrolabe comes out ahead. At a moderate load of twenty queries per second, it cuts mean time-to-first-token by thirty percent compared to the best baseline. At higher loads, the gains are even bigger.

Jane: And it's not just about latency. They also track how many requests get preempted—that's when a server runs out of memory and has to pause a request to make room for others. At their capacity limit, Astrolabe had about three hundred preemptions while the baseline had around one thousand eight hundred. That's a six-fold reduction.

Tom: That's a huge deal for user experience. Preemptions mean wasted work and slower responses. So Astrolabe isn't just making things a little faster; it's making the whole system more stable and predictable.

Jane: And the beauty is that it's all one-shot dispatch. No moving requests around after they start. That simplicity is what keeps the overhead low.

Tom: So the summary is: predict the length, simulate the latency, and pick the best of two random options. Simple in concept, but clearly powerful in practice. Next up, we'll look at some of the deeper improvements and what happens when you push the system to its limits.

Improvements: Tom: Welcome back to our deep dive on "Astrolabe: Balancing Load in LLM Serving with Randomized Prediction-Guided Scheduling." So far we've covered the core idea. Now, Jane, what are some of the clever improvements and details that make this paper stand out?

Jane: One thing I really appreciate, Tom, is that they didn't just stop at the default setting. They ran a whole ablation study on the "k" in power-of-k choices. The default is k=two but they tested k=four k=eight and even k=twelve which is full fanout—checking every server. And the results are fascinating.

Tom: What did they find?

Jane: The capacity—the maximum load they could handle—was basically the same for all values of k. It was around thirty-one point six queries per second. But the scheduling overhead grew with k. So the cheapest option, k=two gives you almost all the benefit for a fraction of the cost. And in some cases, the full fanout actually did *worse* on latency because of herding.

Tom: That's a great example of why the randomness isn't just a nice-to-have; it's a core part of the design. They also tested what happens with a perfect length predictor. They call it the "oracle" version. How much does the prediction accuracy actually matter?

Jane: Surprisingly little, actually. With the oracle, capacity went up by about three percent—from thirty-one point six to thirty-two point four QPS. That tells us that the system is already pretty robust to prediction errors. The simulation and the power-of-two choices are doing the heavy lifting. Even a rough length guess is good enough to get most of the benefit.

Tom: And they stress-tested that robustness even further. They injected artificial noise into the predictions—up to one hundred percent error—and the system still held up. The worst case was about a six percent increase in mean latency. That's a very resilient design.

Jane: They also tested it under bursty traffic, which is how real-world workloads actually behave. People don't send requests at a perfectly steady rate. And Astrolabe handled bursts much better than the baseline, with up to nine percent lower latency under the most bursty conditions.

Tom: And I have to mention the head-to-head against Llumnix with full migration enabled. On a two-node A100 cluster, Astrolabe outperformed it by up to two point six times in throughput. The migration overhead just collapses under high load, while Astrolabe's one-shot dispatch keeps chugging along.

Jane: It's a strong argument that proactive prediction beats reactive rebalancing. And they also showed it works across different models and datasets. They tested Qwen2-7B and a different workload, and Astrolabe still improved capacity by up to six percent.

Tom: So the improvements aren't just incremental tweaks. They're about making the system robust, efficient, and general. Let's wrap up with our final thoughts on what this means for the future.

Conclusion: Tom: Alright, we've reached the end of our journey through "Astrolabe: Balancing Load in LLM Serving with Randomized Prediction-Guided Scheduling." Jane, give us the final takeaway.

Jane: The big picture, Tom, is that Astrolabe shows you can solve load balancing in LLM serving by being smart *before* you act, rather than fixing problems after they happen. By predicting response lengths and simulating latency on just two random candidate servers, it achieves better performance than systems that do expensive live migration.

Tom: And it does it with lower overhead and more stability. We saw up to thirty-six percent lower time-to-first-token, six times fewer preemptions, and it matched or beat the capacity of every baseline they tested. The numbers are compelling.

Jane: The authors, Wei Da and Evangelia Kalyvianaki from Cambridge, have built something that's practical, not just theoretical. They've shown it works across different models, different workloads, and even under injected noise. That's the mark of a well-engineered system.

Tom: And the implications for the real world are significant. Every company running LLM services—chatbots, coding assistants, search—is paying for GPU time. Astrolabe offers a way to serve more users with the same hardware, or serve the same users with better latency. That's a direct win for both cost and user experience.

Jane: And the design is modular. The prediction interface is pluggable, so as better simulators or length estimators come along, they can be swapped in without redesigning the whole system. That's good engineering.

Tom: We should also note the future work they mentioned: larger-scale testing, more failure handling, and adapting the length model as workloads drift. There's plenty of room to build on this foundation.

Jane: So as we say goodbye to Astrolabe, I'm left feeling optimistic. This is the kind of research that quietly makes the AI infrastructure we all rely on just work better. And that benefits everyone.

Tom: Well said, Jane. That's all the time we have for this paper. Thanks for listening, and join us next time when we'll be exploring another exciting development from the world of AI systems research. Until then, take care!

More episodes

← Home