Scheduling Mixed RL Rollouts Beyond Prefix Locality

arXiv:2608.11152 · cs.DC, cs.LG · Submitted 2026-08-11 · Read on arXiv

Zetao Hong, Song Yuan, Yuanhao Ding, Yibo Zhu, Daxin Jiang, Zhibin Wang, Chen Tian

State Key Laboratory of Novel Software Technology, Nanjing University · StepFun

cs.DC, cs.LG

Submitted: 2026-08-11

Updated: 2026-08-12

Code: https://github.com/vllm-project/router

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 81/100

The gist: MISA-T is a routing-layer admission policy for mixed rollout serving in reinforcement learning (RL) post-training pipelines for large language models (LLMs).

Terminology

Summary

MISA-T is a routing-layer admission policy for mixed rollout serving in reinforcement learning (RL) post-training pipelines for large language models (LLMs). It addresses the problem that prefix-aware routing, while improving inference efficiency through cache reuse and load balancing, does not control how heterogeneous rollout sessions (RLVR, RLHF, and agentic) compete for KV-cache capacity. These workloads have distinct sequence structures, interaction patterns, and KV-residency times, creating substantially different serving demands. Rollout scheduling must account for this heterogeneity without distorting the workload mixture specified by the trainer.

MISA-T combines three components: adaptive session admission, workload-aware KV-capacity allocation, and residency-time-aware KV accounting. The paper states: "MISA-T separates the trainer’s mixture contract from the router’s resource controller: the trainer determines the target distribution, while the router performs session-aware admission control, partitions protected KV capacity by workload class, and accounts for class-specific KV residency time."

The contributions are:

  1. Overload-aware session admission: "We derive an adaptive admission cap for new sessions from observed KV demand and overload pressure. Requests beyond the cap are held and periodically re-evaluated, while existing continuations remain protected, limiting KV-cache churn and reducing the need to retune static concurrency limits."

  2. Workload-aware capacity allocation: "We partition protected admission capacity across RLVR, RLHF, and agentic workloads according to their demand and KV footprints. This allows resource control to reflect heterogeneous sequence structures while leaving the target workload composition under trainer control."

  3. Residency-time-aware KV accounting: "We weight each workload’s KV demand by its observed session residency time, capturing both inference time and tool-interaction intervals during which reusable KV remains resident. This extends footprint-based allocation to account for KV block-time demand."

The key insight is that Admission is a KV commitment. Admitting a new session is not an isolated request-placement decision; it commits capacity for subsequent KV growth, future continuations, and intervals between turns. The paper formalizes this with a block-time demand proxy: Rw,b = Nw,b k̄b Tbb, where Nw,b is the number of unfinished sessions, k̄b is the mean KV footprint, and Tbb is the residency time for class b.

In rollout-only ablations on Step3.7 and Qwen3.6-35B-A3B, MISA-T improves rollout throughput over a sweep-tuned cache-aware vLLM Router by 53.3% and 43.6%, respectively, while maintaining high prefix-cache hit rates (97.8% and 95.3%). The paper reports: Relative to sweep-tuned vLLM Router, MISA-T improves rollout throughput by 53.3% on Step3.7 and 43.6% on Qwen3.6-35B-A3B, with prefix-cache hit rates of 97.8% and 95.3%, respectively.

In a matched 50-iteration Step3.7 end-to-end training experiment, MISA-T increases rollout throughput by 35.6%, reduces mean iteration time by 22.8%, and raises prefix-cache hit rate from 74.5% to 96.2%. The paper states: Under the same 50-iteration training horizon, MISA-T increases rollout throughput by 35.6%, reduces mean iteration time by 22.8%, and raises prefix-cache hit rate from 74.5% to 96.2%. The consumed workload mixture stays close to the trainer target: "The target Agent/RLVR/RLHF mixture is 48.6/33.1/18.2%. Over 50 iterations, MISA-T consumes 45.9/33.3/20.8%, with a total-variation distance of 2.71 percentage points, compared with 48.3/29.4/22.4% and 4.14 points under vLLM Router. Task scores remain comparable, with the absolute pass@4 difference between MISA-T and vLLM Router is below 0.5 percentage points on SWE-Pro, SWE-Verified, and SWE-MTLG."

The ablation shows the incremental value of each component: Relative to Session Admission, MISA improves rollout throughput by 10.4% on Step3.7 and 9.7% on Qwen3.6-35B-A3B. Adding residency-time weighting improves it by a further 24.3% and 9.7%, respectively.

The paper also demonstrates compatibility with CPU KV-cache offloading: "After the initial warm-up, Figure 5 shows that GPU KV-cache usage remains close to or above 90% through most of the steady-state interval... CPU KV cache offloading is compatible with MISA-T, allows the GPU KV cache to remain highly utilized, and improves mean per-replica RPM by 35.6%."

The paper concludes: "Mixed RL rollout serving requires more than prefix-aware instance selection. Under high concurrency, the routing layer must control how many sessions compete for KV-cache capacity and allocate protected capacity across workloads with different KV footprints and residency times, while leaving the target workload mixture under trainer control. MISA-T addresses these requirements through adaptive session admission, workload-aware session caps, and residency-time-aware KV accounting."

Limitations noted: "MISA-T assumes that requests carry workload labels and relies on timely serving-state reports from inference instances. Delayed or incomplete snapshots can temporarily reduce the accuracy of KV-demand estimates and admission caps."

Improvements for AI systems

Improvements to AI Systems Based on MISA-T:

  1. Adaptive Admission Control for Multi-Tenant Serving: AI systems can implement overload-aware session admission that dynamically caps new sessions based on real-time KV-cache demand, rather than static concurrency limits. This prevents cache thrashing and maintains throughput under mixed workloads (e.g., RLVR, RLHF, agentic) without manual retuning.

  2. Workload-Class-Aware Resource Partitioning: Systems can allocate protected KV-cache capacity per workload class (e.g., 48% agentic, 33% RLVR, 19% RLHF) based on observed demand and footprint, ensuring fair resource sharing while preserving the trainer’s target mixture. This avoids one workload starving others and keeps training dynamics stable.

  3. Residency-Time-Weighted KV Accounting: AI systems can weight KV demand by session residency time (including tool-interaction pauses), not just instantaneous footprint. This improves capacity planning for interactive/agentic workloads, reducing premature evictions and boosting cache hit rates (e.g., from 74.5% to 96.2%).

  4. Trainer-Router Separation of Concerns: The system can decouple the trainer’s mixture contract from the router’s resource control, allowing the trainer to specify target distributions while the router independently manages admission and capacity. This yields closer adherence to the target mixture (total-variation distance reduced from 4.14 to 2.71 percentage points) without distorting training data.

  5. Higher Rollout Throughput and Lower Iteration Time: By combining the above, AI systems can achieve 35.6% higher rollout throughput and 22.8% lower mean iteration time in end-to-end RL training, enabling faster model iteration and more efficient GPU utilization.

  6. Seamless CPU KV-Cache Offloading Integration: The system can maintain >90% GPU KV-cache utilization during steady state while offloading to CPU, improving per-replica requests per minute by 35.6%—making it viable for memory-constrained deployments.

  7. Graceful Degradation Under Stale State: While relying on timely serving-state reports, the system can be extended with fallback mechanisms (e.g., conservative caps or periodic re-sync) to handle delayed snapshots, reducing temporary admission inaccuracies.

What the improved AI system can do:

  • Serve heterogeneous RL rollout workloads concurrently with high cache efficiency and throughput.

  • Automatically adapt to changing session arrivals and KV footprints without manual tuning.

  • Maintain training fidelity by honoring the trainer’s workload mixture within 2.7% total-variation distance.

  • Support long-horizon agentic tasks with tool-use pauses without wasting GPU memory.

  • Scale to larger models and higher concurrency while keeping task performance (pass@4) within 0.5 percentage points of baseline.

Sources

Related papers