Astrolabe: Balancing Load in LLM Serving with Randomized Prediction-Guided Scheduling
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Astrolabe: Balancing Load in LLM Serving with Randomized Prediction-Guided Scheduling".
Jane: The paper was written by Wei Da and Evangelia Kalyvianaki from University of Cambridge.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Title and Authors: Tom: Welcome back to the show, everyone! I'm Tom, and joining me as always is the brilliant Jane. Today we're cracking open a fresh arXiv paper that's got the systems community buzzing. It's called "Astrolabe: Balancing Load in LLM Serving with Randomized Prediction-Guided Scheduling."
Jane: And I'm Jane! Tom, I have to say, I love this title. An astrolabe was an ancient device used by sailors to navigate by the stars, right? It helped you figure out where you were and where to go. And this paper is doing exactly that for large language model servers—helping them navigate where to send requests.
Tom: That's a perfect way to put it, Jane. The authors are Wei Da and Evangelia Kalyvianaki from the University of Cambridge. And they're tackling a problem that anyone running a big AI service feels every single day. When you have a cluster of GPUs serving chatbots or code generators, how do you decide which GPU gets which user request?
Jane: Right, and the obvious answer seems easy—just spread them out evenly, right? But it's not that simple. The paper explains that LLM inference is just wildly unpredictable. You don't know how long a response will be when a request comes in. One user asks for a haiku, another asks for a full essay. And that uncertainty wreaks havoc on load balancing.
Tom: Exactly. The old-school approach is like round-robin, just sending requests to servers in a circle. But that ignores the fact that some requests are heavy and some are light. It's like a grocery store where every cashier gets the same number of customers, but one customer has a cart overflowing while another has a single loaf of bread. The lines get uneven fast.
Jane: And the more sophisticated systems try to fix this by migrating requests between servers mid-flight. That means moving the "KV cache"—the memory of what the model has already computed—across the network. But that transfer is expensive and can clog up the network, especially when things get busy.
Tom: So Astrolabe's big idea is to do the smart thing *before* sending the request, not after. They predict how long the response will be, they simulate what the latency would be on a couple of candidate servers, and then they send the request to the best one. No migration needed.
Jane: It's proactive instead of reactive. And the "randomized" part in the title is clever too. Instead of checking every server, they just check two random ones and pick the better. It's a classic trick from distributed systems called "power of two choices," and it prevents a stampede where every request tries to go to the same "best" server.
Tom: And the results are pretty stunning. On their test cluster, they matched the capacity of the best existing system but cut the time-to-first-token by up to thirty-six percent and reduced preemptions by six times. We'll get into the nitty-gritty of those numbers in a bit.
Jane: So this isn't just a theoretical idea. They built it, they tested it against real baselines, and it wins. I'm excited to dig into how the prediction actually works.
Tom: Stick around—that's exactly what we're covering next.
Summary: Tom: Welcome back. We're diving into "Astrolabe: Balancing Load in LLM Serving with Randomized Prediction-Guided Scheduling." Last time we set the scene—the problem of unpredictable request lengths and the cost of migration. Now let's talk about the actual machinery. Jane, how does Astrolabe pull off this prediction trick?
Jane: So there are three main pieces working together, Tom. First, there's a "length tagger." It's a small, fast machine-learning model—a fine-tuned RoBERTa—that looks at the user's prompt and guesses how many tokens the response will be. It's not perfect, but it's pretty good, with a mean absolute error of about seventy-nine tokens.
Tom: And that guess feeds into the second piece, the "Predictor" sidecar that runs on each server. This is the really cool part. Each server has a little simulator that knows its own current state—how many requests are running, how much memory is free, what's in the queue. When a new request comes in, the simulator plays it forward against that real state and estimates the end-to-end latency.
Jane: Right, and they built this on top of Vidur, which is an existing LLM serving simulator. But they had to rework it to run online, in real time, instead of offline for planning. They made it fast enough to give a prediction in milliseconds.
Tom: And the third piece is the global scheduler, the "brain" that makes the final call. When a request arrives, the scheduler picks two random servers, asks their Predictors for a latency estimate, and sends the request to whichever one predicts a lower latency. That's the power-of-two choices strategy.
Jane: And it's worth emphasizing why that randomness matters. If the scheduler always picked the absolute best server according to the predictions, then under a burst of traffic, every request would see the same "best" answer and pile onto one server. That's called herding. By sampling randomly, they break that pattern.
Tom: The paper compares this against five other schedulers, including round-robin and the heuristic used by the Llumnix system. And across the board, Astrolabe comes out ahead. At a moderate load of twenty queries per second, it cuts mean time-to-first-token by thirty percent compared to the best baseline. At higher loads, the gains are even bigger.
Jane: And it's not just about latency. They also track how many requests get preempted—that's when a server runs out of memory and has to pause a request to make room for others. At their capacity limit, Astrolabe had about three hundred preemptions while the baseline had around one thousand eight hundred. That's a six-fold reduction.
Tom: That's a huge deal for user experience. Preemptions mean wasted work and slower responses. So Astrolabe isn't just making things a little faster; it's making the whole system more stable and predictable.
Jane: And the beauty is that it's all one-shot dispatch. No moving requests around after they start. That simplicity is what keeps the overhead low.
Tom: So the summary is: predict the length, simulate the latency, and pick the best of two random options. Simple in concept, but clearly powerful in practice. Next up, we'll look at some of the deeper improvements and what happens when you push the system to its limits.
Improvements: Tom: Welcome back to our deep dive on "Astrolabe: Balancing Load in LLM Serving with Randomized Prediction-Guided Scheduling." So far we've covered the core idea. Now, Jane, what are some of the clever improvements and details that make this paper stand out?
Jane: One thing I really appreciate, Tom, is that they didn't just stop at the default setting. They ran a whole ablation study on the "k" in power-of-k choices. The default is k=two but they tested k=four k=eight and even k=twelve which is full fanout—checking every server. And the results are fascinating.
Tom: What did they find?
Jane: The capacity—the maximum load they could handle—was basically the same for all values of k. It was around thirty-one point six queries per second. But the scheduling overhead grew with k. So the cheapest option, k=two gives you almost all the benefit for a fraction of the cost. And in some cases, the full fanout actually did *worse* on latency because of herding.
Tom: That's a great example of why the randomness isn't just a nice-to-have; it's a core part of the design. They also tested what happens with a perfect length predictor. They call it the "oracle" version. How much does the prediction accuracy actually matter?
Jane: Surprisingly little, actually. With the oracle, capacity went up by about three percent—from thirty-one point six to thirty-two point four QPS. That tells us that the system is already pretty robust to prediction errors. The simulation and the power-of-two choices are doing the heavy lifting. Even a rough length guess is good enough to get most of the benefit.
Tom: And they stress-tested that robustness even further. They injected artificial noise into the predictions—up to one hundred percent error—and the system still held up. The worst case was about a six percent increase in mean latency. That's a very resilient design.
Jane: They also tested it under bursty traffic, which is how real-world workloads actually behave. People don't send requests at a perfectly steady rate. And Astrolabe handled bursts much better than the baseline, with up to nine percent lower latency under the most bursty conditions.
Tom: And I have to mention the head-to-head against Llumnix with full migration enabled. On a two-node A100 cluster, Astrolabe outperformed it by up to two point six times in throughput. The migration overhead just collapses under high load, while Astrolabe's one-shot dispatch keeps chugging along.
Jane: It's a strong argument that proactive prediction beats reactive rebalancing. And they also showed it works across different models and datasets. They tested Qwen2-7B and a different workload, and Astrolabe still improved capacity by up to six percent.
Tom: So the improvements aren't just incremental tweaks. They're about making the system robust, efficient, and general. Let's wrap up with our final thoughts on what this means for the future.
Conclusion: Tom: Alright, we've reached the end of our journey through "Astrolabe: Balancing Load in LLM Serving with Randomized Prediction-Guided Scheduling." Jane, give us the final takeaway.
Jane: The big picture, Tom, is that Astrolabe shows you can solve load balancing in LLM serving by being smart *before* you act, rather than fixing problems after they happen. By predicting response lengths and simulating latency on just two random candidate servers, it achieves better performance than systems that do expensive live migration.
Tom: And it does it with lower overhead and more stability. We saw up to thirty-six percent lower time-to-first-token, six times fewer preemptions, and it matched or beat the capacity of every baseline they tested. The numbers are compelling.
Jane: The authors, Wei Da and Evangelia Kalyvianaki from Cambridge, have built something that's practical, not just theoretical. They've shown it works across different models, different workloads, and even under injected noise. That's the mark of a well-engineered system.
Tom: And the implications for the real world are significant. Every company running LLM services—chatbots, coding assistants, search—is paying for GPU time. Astrolabe offers a way to serve more users with the same hardware, or serve the same users with better latency. That's a direct win for both cost and user experience.
Jane: And the design is modular. The prediction interface is pluggable, so as better simulators or length estimators come along, they can be swapped in without redesigning the whole system. That's good engineering.
Tom: We should also note the future work they mentioned: larger-scale testing, more failure handling, and adapting the length model as workloads drift. There's plenty of room to build on this foundation.
Jane: So as we say goodbye to Astrolabe, I'm left feeling optimistic. This is the kind of research that quietly makes the AI infrastructure we all rely on just work better. And that benefits everyone.
Tom: Well said, Jane. That's all the time we have for this paper. Thanks for listening, and join us next time when we'll be exploring another exciting development from the world of AI systems research. Until then, take care!
Wei Da, Evangelia Kalyvianaki
University of Cambridge
cs.DC, cs.AI
Submitted: 2026-08-11
Updated: 2026-08-12
Comments: 16 pages. Accepted at SYSTOR 2026. Camera-ready version with expanded evaluation and revisions. Previously circulated as "Block"; renamed "Astrolabe" to match the SYSTOR publication title
Code: https://github.com/AKafakA/Block
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 52/100
The gist: This paper presents Astrolabe, a randomized prediction-guided scheduler for one-shot request dispatch in multi-instance large language model (LLM) serving.
Terminology
Summary
This paper presents Astrolabe, a randomized prediction-guided scheduler for one-shot request dispatch in multi-instance large language model (LLM) serving. Astrolabe improves load balancing without relying on migration-based rebalancing, whose KV-cache transfers can introduce substantial overhead and network contention under high load. Astrolabe combines response-length estimation, per-instance simulation-based latency prediction, and a power-of-two choices dispatch policy to improve load balancing and avoid request herding. On the default Llama-2-7B/ShareGPT setup, Astrolabe matches the best load-aware baseline's SLO capacity (31.6 vs 31.5 QPS) while improving latency metrics: 8–36% lower mean time-to-first-token (TTFT), 16–77% lower P99 TTFT, up to 5.6% lower mean end-to-end (E2E) latency, and 6× fewer preemptions once capacity is reached. Under configuration shifts, Astrolabe also improves SLO capacity by up to +6% on Qwen2-7B and +7.1% under tight batching, with 6–9% lower mean E2E under bursty arrivals, while achieving a 2.8-fold reduction in per-predictor CPU relative to full fanout. With migration enabled on A100, Astrolabe outperforms Llumnix by up to 2.6× in throughput with orders-of-magnitude lower per-token latency at saturation.
The paper notes that "the rise of large language models like GPT [1], Llama [62], Gemini [61], Qwen [72], and DeepSeek [41] has revolutionized modern applications, such as chatbots [1], virtual assistants [31], code generation [34], and creative writing [32]." This places pressure on LLM inference serving systems to meet latency requirements and maximize throughput. Key techniques such as continuous batching [74], Paged Attention [39], chunked prefill [4], and FlashAttention [19] have been proposed.
However, "LLM inference is often characterized as unpredictable [60]. For example, LLMs generate tokens autoregressively based on preceding tokens until a stop signal is reached, leading to variable response lengths and numbers of decoding steps [57]. Paged Attention [39], which dynamically allocates memory resources and allows request preemption, further contributes to the dynamic nature of runtime memory consumption. Decode-step latency also exhibits high variance due to batch-size variation and interference between requests in the same batch [60]."
These uncertainties make load balancing across LLM serving instances challenging: "common runtime metrics (latency, throughput, memory) no longer represent each request's E2E load. Moreover, unexpectedly long responses further load already-overloaded hosts and block subsequent requests (head-of-line blocking [16, 51]). In practice, production multi-instance frameworks [6, 47] dispatch with simple heuristics like round-robin, offering no performance guarantee."
Alternative solutions like Llumnix [60] combine dispatching with dynamic rebalancing through live migration, but "migration requires transferring request KV caches [38] across instances, adding network, memory, and orchestration overhead. Under high load, these costs can accumulate and contend with normal serving traffic, potentially degrading tail latency and throughput."
The paper identifies an opportunity: if length and latency can be estimated at dispatch time, the scheduler can proactively route long requests to less-loaded instances before imbalance manifests.
A basic approach of querying every instance's predictor in parallel (full fanout) has three costs: "First, each dispatch costs O(N) predictor queries for N instances, which scales poorly in the scheduler control plane. Second, simulation runs on CPU to predict the scheduling signal (e.g., latency) for each candidate instance, so aggregate cluster-CPU grows linearly with N per request. Third, deterministic fanout creates a thundering-herd issue: under short correlated bursts, nearby requests compute similar 'best' candidates and route to the same instance, overloading it transiently [8, 44]."
Astrolabe comprises four services: a query length tagger, a global scheduler, per-instance Predictor sidecars, and inference framework backends.
The query length tagger is the entry point and predicts the anticipated response length for each request. Requests are then forwarded to the global scheduler, which selects a target model instance. For each incoming request, the global scheduler issues predict RPCs to a sampled subset of Predictor sidecars (power-of-k choices; k=2 by default, and k=N recovers full fanout). Each Predictor queries its colocated backend via a status API to obtain the instance's current runtime state (e.g., running/waiting requests, batching state, and memory availability), and then runs a lightweight simulation to estimate the request's target metric (e.g., E2E latency). The global scheduler picks the target instance from these predictions per the configured policy, and the backend executes the request and returns the response.
A model instance is the collection of services on a GPU host driving one or more GPUs: the inference framework itself plus a sidecar service called the Predictor.
The Predictor's main role is to predict key performance metrics, such as E2E latency or TTFT, for incoming requests. "The Predictor runs locally on each instance, aggregates runtime data from the status API, and serves a scalar metric prediction via the predict API to the global scheduler. In cases where a request's decoded length exceeds its predicted length at runtime, the simulator extends the current prediction by assuming the request will complete within the next 10 tokens by default."
The prototype integrates Vidur [2], redesigned for single-instance prediction and encapsulated behind the Predictor API.
Simulation proceeds in two stages: a local-scheduler simulator models the backend's batching strategy, then a linear model predicts execution times for the generated batches.
The global scheduler service is stateless and lightweight without maintaining a global table of instance statuses.
Upon each request arrival, the scheduler issues a predict RPC to sampled Predictor sidecars and then dispatches the request to the instance with the lowest predicted metric.
E2E latency is the default target metric.
"This differs from centralized predictors that need a global per-instance status view (e.g., the estimation-driven global schedulers in TetriServe [33] and DynamoLLM [59]). In Astrolabe, prediction runs on each instance in parallel using its local state, and each Predictor returns only a scalar metric (e.g., predicted latency). Decision time is bounded by the slowest Predictor reply, while avoiding the cost of exporting heavy instance state (e.g., request lists) to a central simulator."
"To avoid the O(N) per-dispatch messaging cost of full fanout and the scheduler-herding issue, Astrolabe's default samples k=2 random instances (power-of-two choices) and queries only those Predictors. Setting k=N recovers full fanout."
"The query length tagger is an online service that estimates the response length of each request from its prompt and the target serving model. It runs in parallel with request handling and uses a lightweight proxy model to keep overhead low." The predicted length is used only for metric prediction and scheduling decisions and does not change the inference parameters and results.
The prototype integrates with vLLM 0.7.2 and comprises about 4,000 LoC. The Predictor internally invokes a backendspecific simulator
based on Vidur's offline simulator codebase, substantially adapted for online, per-request use.
The authors rewrote the primary simulation hot paths and added a per-Predictor cache that memoizes latency predictions per (batch size, token count), substantially reducing the per-request simulation cost.
For the query length tagger, the authors fine-tuned a lightweight RoBERTa-base [42] regression model (125M parameters).
The length estimation model achieves MAE 78.8 and MAPE 24.4% on the 10k test split, which is comparable to the 7B prompt-based model reported in Sequence Scheduling [77].
The primary cluster consists of 12 d7525 nodes on CloudLab [22], each with two 16-core AMD 7302 CPUs, 128 GB ECC memory, one NVIDIA A30 GPU (24 GB), and 25 Gbps NICs. A second 2-node CloudLab d8545 cluster has 4× A100-40GB GPUs per node and 100 Gbps NICs. The main experiments use ShareGPT [71] with Llama-2-7B [62] in FP16 precision. The model weights occupy 12.5 GB of GPU memory, split into 1056 vLLM memory blocks for KV cache. The maximum request batch size is 48, and the Chunked-Prefill chunk size is 512.
The length prediction model reaches MAE 78.8 and MAPE 24.4% on the 10k test split.
For simulation-based metric prediction, Enabling Chunked Prefill yields much lower error than vLLM's default Prioritized Prefill.
The scheduler maintains a high probability (40 to 80%) of selecting the Rank-1 instance
among candidates. For Llama-2-70B on the A100 cluster with TP=4, the simulator achieves a latency-prediction error of 16.6%.
Across the 17-point QPS sweep, "Astrolabe's peak reductions over the best load-aware baseline are 38% mean TTFT, 78% TTFT-P99, 7% mean E2E, and 12% E2E-P99, with up to 5% higher throughput; relative to the worst baseline, the gains reach 97% lower TTFT and 53% higher throughput." Astrolabe's per-request overhead stays under 3% of E2E within capacity.
At QPS=32, Llumnix- accumulates 1.8k preemptions vs Astrolabe's 300 (6× fewer).
The serving capacity is 31.6 QPS for Astrolabe vs 31.5 for Llumnix-.
On the A100 cluster with full migration enabled, Llumnix collapses at QPS=20, where mean request latency jumps from 6.3 s to 106 s and mean token latency from 21 ms to 1.15 s.
At QPS=36, "Astrolabe without CP delivers 2.57× the throughput (18.7k vs 7.3k tok/s), 36× lower mean request latency (8.2 s vs 300 s), and 123× lower mean token latency (26 ms vs 3.2 s) than Llumnix. Enabling CP widens these to 2.58×/40×/135×."
Capacity is effectively tied from k=2 upward: Po2-est reaches 31.6 QPS vs Po4-est 31.9 (the maximum), Po8-est 31.7, and Fanout-est 31.7, all within a ±0.3 QPS band at SLO = 10s.
The paper notes that k=4 gives the best empirical tradeoff: it has the highest capacity and throughput, while retaining the second-lowest E2E latency and scheduling overhead.
Oracle adds 0.8 QPS for Po2 (31.6 → 32.4) and 0.9 QPS for Fanout (31.7 → 32.6) at SLO = 10s, a 3% capacity gain.
Under high burstiness (α=0.25), Astrolabe's mean E2E latency is about 9% lower than Llumnix-; under near-regular arrivals (α=2.0), the gap narrows but Astrolabe retains a small edge.
Moving from α=2.0 to α=0.25 raises mean E2E by 13% for Astrolabe vs 22% for Llumnix-.
The largest mean degradation is +6.3% at (100%, 25%), while the extreme (100%, 100%) cell is +5.1%; most cells remain within the ±4% single-deploy noise floor.
The paper notes that P99 degradation remains below Llumnix- in every cell (maximum +3.7%).
At QPS=20 mean per-predictor CPU is 20.3% for Po2 and 56.3% for Fanout, a 2.8-fold reduction.
Per-predictor RSS is Po2 1225 MB, Fanout 1773 MB at QPS=20.
Across configuration variations: with batch size reduced 48→24, Astrolabe reaches 28.5 QPS vs Llumnix- 26.6 QPS (+7.1%)
; with Qwen2-7B, Astrolabe reaches 73.9 QPS vs Llumnix- 69.7 QPS (+6.0%, +12.6% with oracle lengths)
; with BurstGPT [67], Astrolabe-oracle reaches 64.5 QPS vs Llumnix- 61.6 QPS (+4.7%).
"Astrolabe combines response-length estimation, simulation-based latency prediction, and randomized dispatch for multi-instance LLM serving. Across our experiments, prediction-guided one-shot dispatch matches or outperforms migration-based rebalancing while avoiding KV-cache transfer overhead. The paper leaves
larger-scale multi-GPU evaluation, richer failure handling, and length-model retraining under workload drift as future work."
Improvements for AI systems
Based on the paper, here are the specific improvements I can implement in an AI system:
-
Add a lightweight RoBERTa-base regression model (125M parameters) that predicts the expected response length from the input prompt before generation begins
-
Use this prediction to inform scheduling decisions, not to alter generation behavior
-
Achieve MAE of 78.8 tokens and MAPE of 24.4% on ShareGPT-style prompts
-
The model runs in parallel with request handling, adding only 4.8ms per request
-
Integrate a simulator (based on Vidur's approach) that predicts end-to-end latency for a request on each candidate instance
-
The simulator accounts for:
-
Current batch composition and size
-
Pending queue depth
-
Available KV cache memory blocks
-
The backend's batching strategy (e.g., chunked prefill vs. prioritized prefill)
-
Predicted response length
-
Maintain a cache that memoizes latency predictions per (batch size, token count) to reduce per-request simulation cost
-
Achieve 10-15% absolute prediction error, with 40-80% probability of selecting the true Rank-1 instance
-
For each incoming request, sample exactly 2 random instances (not all instances)
-
Query both instances' predictors in parallel
-
Route the request to the instance with the lower predicted latency
-
This reduces per-request predictor fanout from O(N) to O(2), cutting per-predictor CPU by 2.8x compared to full fanout
-
Randomization breaks herding: simultaneous requests are unlikely to sample the same pair, preventing transient overload
-
Implement the auto-extension rule: if a request's actual length exceeds prediction, extend the simulation assuming completion within the next 10 tokens
-
Use relative ranking (not absolute latency values) for instance selection, so common-mode errors cancel out
-
The system tolerates up to ±100% independent Gaussian noise on latency predictions with at most +6.3% mean E2E degradation and +3.7% P99 E2E degradation
-
Default to E2E latency as the scheduling objective, but keep the interface configurable
-
Support alternative objectives (e.g., TTFT, throughput) via a pluggable interface
-
Allow tuning the number of sampled instances (k) from 2 to N, where k=N recovers full fanout
For a multi-instance LLM serving cluster (e.g., 12 GPUs):
-
Achieve 31.6 QPS SLO capacity (P99 TTFT ≤ 10s) on Llama-2-7B/ShareGPT, matching the best load-aware baseline (31.5 QPS) and full-fanout prediction (31.7 QPS) while using 6x fewer predictor queries
-
Reduce latency metrics:
-
8-36% lower mean TTFT vs. best baseline
-
16-77% lower P99 TTFT
-
Up to 5.6% lower mean E2E latency
-
Up to 97% lower mean TTFT vs. round-robin
-
Reduce preemptions by 6x at overload (300 vs. 1,800 preemptions at QPS=32), avoiding KV-cache recomputation stalls
-
Handle configuration shifts automatically:
-
+7.1% capacity gain with tight batching (batch size 48→24)
-
+6.0% capacity gain on Qwen2-7B
-
+4.7% capacity gain on BurstGPT dataset (oracle lengths)
-
Maintain performance under bursty arrivals: 6-9% lower mean E2E latency vs. Llumnix- across gamma-distributed arrival patterns (shape parameter 0.25 to 2.0)
-
Outperform migration-based rebalancing (Llumnix) on A100 clusters:
-
Up to 2.6x higher throughput at saturation
-
Orders-of-magnitude lower per-token latency (26ms vs. 3.2s at QPS=36)
-
Avoids KV-cache transfer overhead and network contention
-
Keep control-plane overhead low: per-request scheduling overhead stays under 3% of E2E latency within capacity, with per-predictor CPU at 20% (vs. 56% for full fanout)
-
Scale to larger models: validated on Llama-2-70B with tensor parallelism (TP=4), achieving 16.6% latency-prediction error
In summary: The improved AI system can serve LLM requests across a cluster with near-optimal load balancing, avoiding the overhead of live migration while matching or exceeding its performance, and remaining robust to prediction errors and workload shifts.
Abstract
This paper presents Astrolabe, a randomized prediction-guided scheduler for one-shot request dispatch in multi-instance large language model (LLM) serving. Astrolabe improves load balancing without relying on migration-based rebalancing, whose KV-cache transfers can introduce substantial overhead and network contention under high load. It combines response-length estimation, per-instance simulation-based latency prediction, and a power-of-two-choices dispatch policy to balance load while avoiding request herding. On the default Llama-2-7B/ShareGPT setup, Astrolabe matches the SLO capacity of the best load-aware baseline (31.6 versus 31.5 QPS), while reducing mean time-to-first-token (TTFT) by 8 to 36 percent, P99 TTFT by 16 to 77 percent, and mean end-to-end (E2E) latency by up to 5.6 percent, with approximately six times fewer preemptions once capacity is reached. Under configuration shifts, Astrolabe improves SLO capacity by up to 6 percent on Qwen2-7B and 7.1 percent under tight batching, reduces mean E2E latency by 6 to 9 percent under bursty arrivals, and achieves an approximately 2.8-fold reduction in per-predictor CPU usage relative to full fanout. With migration enabled on A100 GPUs, Astrolabe outperforms Llumnix by up to 2.6 times in throughput while achieving orders-of-magnitude lower per-token latency at saturation.
Sources
- GPT-4 Technical Report
- Vidur: A Large-Scale Simulation Framework For LLM Inference
- Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
- SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
- Revati: Transparent GPU-Free Time-Warp Emulation for LLM Serving
- Predictive Scheduling for Efficient Inference-Time Reasoning in Large Language Models
- SLOs-Serve: Optimized Serving of Multi-SLO LLMs
- SCOOT: SLO-Oriented Performance Tuning for LLM Inference Engines
- ELIS: Efficient LLM Iterative Scheduling System with Response Length Predictor
- Dodoor: Efficient Randomized Decentralized Scheduling with Load Caching for Heterogeneous Tasks and Clusters
- LLM-Emu: Native Runtime Emulation of LLM Inference via Profile-Driven Sampling
- FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
- BEST-Route: Adaptive LLM Routing with Test-Time Optimal Compute
- From LLM to NMT: Advancing Low-Resource Machine Translation with Claude
- TurboTransformers: An Efficient GPU Serving System For Transformer Models
- ServerlessLLM: Low-Latency Serverless Inference for Large Language Models
- DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU Multiplexing
- JSONSchemaBench: A Rigorous Benchmark of Structured Outputs for Language Models
- Intelligent Virtual Assistants with LLM-based Process Automation
- A Confederacy of Models: a Comprehensive Evaluation of LLMs on Creative Writing
Related papers
- iScheduler: Reinforcement Learning-Driven Continual Optimization for Large-Scale Resource Investment Problems
- SAMM: Sharded Automated Market Maker
- InferScale: GPU-Native KV Injection for Personalized LLM Serving
- Vigil: Accountable Liveness against Selective Silence
- Steelhead: Interleaving Partially Synchronous and Asynchronous Commit Rules on a Shared DAG
- Pushing CPU Speech Synthesis to the Wall: Extreme Inference Tuning under Serverless Architecture and Billing