TideRL: Boosting Agentic RL Goodput with Readiness-Aware Scheduling

arXiv:2608.10402 · cs.LG, cs.DC · Submitted 2026-08-11 · Read on arXiv

Yanyu Ren, Xizheng Wang, Xiao Liu, Bowen Lv, Hanchen Zhang, Shudan Zhang, Hanyu Lai, Shuai Wang, Li Chen, Dan Li, Jie Tang

Tsinghua University · Z.AI · Zhongguancun Laboratory

cs.LG, cs.DC

Submitted: 2026-08-11

Updated: 2026-08-12

Code: https://github.com/THUDM/slime

License: http://creativecommons.org/licenses/by-nc-nd/4.0/

Importance score: 75/100

The gist: TIDE RL: Boosting Agentic RL Goodput with Readiness-Aware Scheduling Abstract Reinforcement learning (RL) for large language models is moving toward multi-turn agentic workloads, where rollout tasks

Terminology

Summary

TIDE RL: Boosting Agentic RL Goodput with Readiness-Aware Scheduling

Abstract

Reinforcement learning (RL) for large language models is moving toward multi-turn agentic workloads, where rollout tasks repeatedly pause for external environments, resume with growing contexts, and finish at highly variable times. In this setting, RL training goodput, measured by training throughput, matters more than raw GPU occupancy: GPU waiting and repeated prefill recomputation are pure overhead. We present TIDE RL, a readiness-aware elastic RL system with Continuous Task Batching, Resource-Aware Ref-Actor Pipelining, and Elastic Resource Scaling. CTB preserves useful rollout state, RA2P selects between decoupled streaming and colocated aggregation from the ready backlog and arrival interval, and ERS moves ranks between rollout and training using the same readiness signals. Across text-only and multi-modal agentic workloads, TIDE RL improves RL training goodput by up to 5.6× over synchronous baselines and over 33% over asynchronous baselines, while reaching similar task performance. It also improves KV cache hit rate by 1.58×, reduces per-step training time by up to 44.3%, and cuts total waiting time by up to 77.6%.

Introduction

Large Language Models (LLMs) are rapidly evolving from static text generators into autonomous agents capable of solving complex, real-world problems. In these agentic workloads, models do not merely produce a single-turn answer; instead, they iteratively interact with external, distributed environments, such as web browsers, smartphones, or computer OSes. This introduces a multi-turn execution flow. An agent generates an action, transmits it over the network to the environment, waits for the execution results, and resumes generation based on the newly returned observations. Because LLM outputs are non-deterministic, execution trajectories of the same task may differ drastically in turn count, context length, and wall-clock duration, producing workloads with extreme variance across both spatial and temporal dimensions.

To enhance these autonomous capabilities, researchers turn to reinforcement learning (RL). While RL has shown immense success in single-turn reasoning tasks, applying it to multi-turn agentic scenarios requires orchestrating massive, networked machine learning clusters. A typical RL workflow goes through four phases. In the rollout phase, a batch of agentic tasks runs concurrently, requesting actions from rollout workers and observations from environment workers to generate interactive trajectories in multiple turns. Later, the RL system evaluates these trajectories in the reward and reference phase. Finally, the system uses the output from the reward and the reference to update the model weights in the actor phase.

However, a small number of long-tail tasks in the rollout phase may block the pipeline and leave expensive GPU resources underutilized. For agentic RL, the right objective is goodput, which we measure through training throughput: the rollout'd tokens that reach the Trainer and advance an update counts. GPU waiting and repeated prefill recomputation are pure overhead. To maximize goodput, asynchronous systems separate these components onto different GPUs, unlike synchronous systems like VeRL that force rollout-train barriers. On a subset of GPUs, rollout worker groups perform continuous generation as producers. On the remaining GPUs, trainer worker groups, which colocate the reference and actor models, pull data from the global buffer to compute weight updates, as consumers. This decoupling theoretically allows generation and training to proceed concurrently, so that the rollout phase of long-tail tasks can take place simultaneously with the training phase of shorter tasks.

While effective for single-turn tasks, current asynchronous architectures exhibit severe systemic inefficiencies when processing multi-turn agentic workloads characterized by high concurrency and variable context lengths. We identify three critical challenges:

  • KV cache preemption during intermittent interactions. Standard inference engines rely on request-level continuous batching, assuming each request is self-contained. However, multi-turn tasks proceed as a sequence of dependent requests separated by idle periods while waiting for external feedback. During these pauses, the engine greedily evicts its KV cache (the memory footprint of historical tokens) to serve requests from other concurrent tasks. Upon resumption, the system suffers a catastrophic cache miss, forcing redundant and computationally expensive recomputation (prefill) for an increasingly long historical context.

  • Stalls and thrashing in reference inference. Existing systems share GPU resources between actor and reference workers within the trainer, enforcing strict stage dependencies. To avoid memory exhaustion, they either wait for large global batches to arrive before initiating updates, or frequently alternate between reference and actor models (which share the same architecture but have different weights) on the same GPUs for each micro-batch. The former creates massive computation bubbles that leave GPUs idle for up to 81% of the step time, while the latter introduces severe model-swapping thrashing overhead.

  • Rigidity of resource allocation under shifting bottlenecks. Agentic RL operates as a producer-consumer pipeline. Because task complexity varies drastically, the system bottleneck dynamically shifts. A batch of complex tasks stalls the trainer as it waits for slow rollout data, while simple tasks create a large ready backlog at the beginning of a training step and then keep feeding new micro-batches rapidly. Relying on static resource partitioning between rollout and trainer worker groups, existing systems cannot react to either the initial ready backlog or the steady ready interval, structurally capping global throughput during these macroscopic workload shifts.

These intertwined challenges expose architectural flaws of existing systems, raising a critical research question: How can we build a unified, elastic training architecture that maximizes RL goodput, as measured by training throughput, under multi-turn agentic execution? The key is not to optimize rollout, training, and resource allocation as isolated modules. Multi-turn RL needs a shared readiness view: Rollout must preserve the task state that will become useful soon, Trainer must choose a ref-actor execution mode that matches how ready data arrives, and the cluster scheduler must move GPUs toward the side that removes the current goodput loss. Answering this requires co-designing the multi-turn memory management, distributed execution pipeline, and dynamic cluster resource allocation.

To this end, we identify two architectural blind spots in existing frameworks. First, systems such as AReaL and StreamRL expose asynchrony, but still treat the rollout-trainer split as a mostly fixed operating point; they cannot use the current ready backlog and arrival pace to jointly choose the ref-actor execution strategy and the GPU split. Second, standard inference engines allocate memory at the request level. Memory management must be elevated to the task level to gain global visibility over the multi-turn lifecycle, actively coordinating context pinning across distributed rollout ranks. Together, these changes make individual data parallel (DP) ranks schedulable resources rather than static members of a worker group.

Motivated by these insights, we present TIDE RL, an elastic and asynchronous distributed system tailored for multi-turn agentic RL goodput. TIDE RL solves these challenges through three co-designed components:

  • Continuous Task Batching (CTB) replaces request-level continuous batching on each rollout rank with task-level admission, preemption, and resumption. CTB tracks per-rank token footprints in real time and applies a semantic-aware priority hierarchy—prioritizing evaluation boundaries, near-complete GRPO groups, active trajectories, and long contexts. This keeps useful rollout state resident until it can feed the global buffer, providing the stable producer side that RA2P and ERS rely on.

  • Resource-Aware Ref-Actor Pipelining (RA2P) resolves the stall-thrashing dilemma through two complementary execution modes selected by two data-readiness signals: Ready-at-Start (RAS), the amount of micro-batch work already present when a training step begins, and Time Per Ready Micro-batch (TPRM), the interval at which additional micro-batches become available. In decoupled mode, Ref and Actor run on separate GPUs in a streaming micro-batch pipeline; a restructured computation graph defers loss computation into the backward pass, enabling Ref forward, Actor forward, and Actor backward to overlap with zero swap overhead. In colocated mode, RA2P uses ready-batch aggregation and zero-copy shared-memory transfer to amortize unavoidable swapping across ready bursts. Together, RA2P turns readiness into goodput under both startup backlog pressure and steady per-micro-batch arrival pressure.

  • Elastic Resource Scaling (ERS) exploits RA2P's dual-mode structure to migrate individual ranks between rollout and training on the fly. ERS reads RAS and TPRM from the global buffer: a small RAS with slow TPRM indicates rollout starvation and triggers colocated RA2P plus more Rollout ranks, whereas a large RAS or fast TPRM indicates trainer-side pressure and triggers decoupled RA2P plus more Trainer capacity. Crucially, ERS piggybacks weight migration onto the natural cache-flush boundary that on-policy RL already mandates at every sync, achieving zero-overhead elasticity: no pipeline suspension, no NCCL group rebuild, no wasted KV cache.

We implement TIDE RL with a highly modular architecture, ensuring seamless compatibility with diverse multi-modal and text-only tasks, and easy integration with mainstream serving and training backends. We deploy the system on a 4-node cluster with 32 NVIDIA H100 GPUs and evaluate models with different architectures. Our evaluations demonstrate large performance gains over existing RL frameworks. End-to-end, TIDE RL raises RL training goodput, measured by training throughput, by 1.8–7.0× on text-only tasks (reducing time-to-convergence by 51.1%) and improves multi-modal goodput by over 33%, reducing training time by 62.2% with similar performance. Component studies close the loop from motivation to design: CTB improves generation goodput by preserving task state, RA2P reduces per-step training time by up to 44.3% under different readiness patterns, and ERS cuts total idle waiting time by 87.4% by scheduling RA2P modes and GPU ranks from the same readiness signals.

Background and Motivation

Agentic RL workloads in this paper follow Group Relative Policy Optimization (GRPO), whose execution alternates between trajectory generation and model update. The agentic GRPO pipeline consists of four sequential stages to train a new model A from the original model O:

  • Rollout Generation: Rollout hosts A to generate actions. Crucially, this involves multi-turn interactions: the model generates an action, pauses to wait for environment execution, receives the observation, and generates the next action. For a given task, a group of G distinct multi-turn trajectories is sampled. Each trajectory may comprise different turns and lengths of messages.

  • Reward Calculation: Once the trajectories in a group are completed, the environment evaluates them. GRPO determines the advantage Ai by standardizing the rewards ri strictly based on the performance of other trajectories within the same group: Ai =(ri −mean(r))/std(r).

  • Reference Calculation: Ref computes the base logprobabilities on O of generated trajectories. This provides a Kullback-Leibler (KL) divergence penalty to prevent the policy from over-optimizing the reward.

  • Actor Update: Actor executes the forward and backward passes on A to update its weights (i.e., update the policy). Once updated, the latest weights are synchronized back to Rollout for the next iteration.

The Trainer comprises Ref, Actor and the reward calculation, and consumes the trajectories from Rollout. Traditional synchronous architectures enforce a strict alternation between the rollout and training stages. This coupling causes severe GPU idling, as the training must wait for the long-tail task to complete. To break this resource coupling, modern systems typically adopt an asynchronous execution paradigm. They decouple generation and training into physically independent workers. This allows Rollout to continuously generate new groups of responses while Trainer consumes historical data in the background, significantly boosting hardware throughput. However, this asynchrony introduces policy staleness, as Rollout generates data using older model weights. Since algorithms like GRPO are fundamentally on-policy, excessive off-policy data severely degrade training stability. Consequently, asynchronous systems must strictly bound this staleness gap, requiring timely weight synchronization to prevent performance degradation.

Key Challenges in Asynchronous Agentic RL

Although asynchronous RL frameworks overlap generation and training, their execution traces under multi-turn agentic workloads reveal three ways the system loses RL goodput.

C1: Catastrophic KV Cache Preemption in Rollouts. Monitoring the KV cache hit rate and generation throughput during a single training step reveals a typical three-phase fluctuation rather than stable cache utilization. At the beginning of the step, the hit rate increases as a large influx of new short sequences is initially cached. However, the hit rate subsequently plummets and remains at only ∼12%, dragging the throughput down to ∼650 tokens/s·instance. A spike in the hit rate occurs at the very end of the training step. The fundamental flaw lies in the mismatch between request-level scheduling and multi-turn semantics. When waiting for environment feedback, the underlying inference engine, driven by standard continuous batching to maximize concurrency, schedules new requests, thereby preempting the tasks' KV caches. The engine suffers massive cache misses as the observation returns, and must recompute the rapidly growing historical context. The late recovery in hit rate is merely an artifact of dropping concurrency: as the global batch nears completion, no new tasks are admitted, allowing the few straggling tasks to safely retain their cache without preemption.

C2: Stalls and Thrashing in Training. Examining the Trainer's execution trace reveals a dilemma between computation stalls and I/O thrashing. Standard execution forces the Trainer to stall significantly while waiting for a full global batch to complete, a delay severely exacerbated by long-tail tasks. Conversely, if we attempt to mitigate this idle time by streaming micro-batches, the system falls into another trap: massive I/O thrashing that degrades training GPU goodput. This dilemma is structurally rooted in the memory-intensive nature of on-policy algorithms like GRPO. The Trainer must execute reference calculations and actor updates sequentially using two distinct models. Under standard execution, the system waits for all rollout tasks in a global batch to finish, making the Trainer's idle time bottlenecked by the slowest long-tail samples. To avoid these macroscopic stalls, the system could adopt a streaming mode to process micro-batches as soon as they arrive. However, existing systems (e.g., VeRL and AsyncFlow) cannot run both models on the same GPUs simultaneously due to the runtime memory consumed by training activations. Consequently, streaming micro-batches forces the system to frequently load and offload model weights between the CPU and GPU over PCIe, replacing the stall with a model-swapping bottleneck.

C3: Shifting Bottlenecks Invalidating Static Allocation. Tracking the Rollout Wait Time (RWT) and Training Wait Time (TWT) across 100 steps under various static Rollout GPU ratios reveals that the system bottleneck oscillates wildly. A lower rollout ratio (0.125) consistently starves the Trainer (high TWT), while a higher ratio (0.5) generates data too fast, causing Rollout to stall (high RWT). Most crucially, even within a single optimal fixed ratio (e.g., 0.25), RWT and TWT frequently cross paths. This invalidation of static partitioning stems from the extreme variance in agentic trajectory execution times, further compounded by the unpredictable load of intermediate evaluation tasks. Unlike single-turn chat, multi-turn environments cause the computational bottleneck to dynamically flip between generation and training, so a static GPU quota cannot preserve goodput across steps. It cannot react to either the micro-batch backlog already ready at a training boundary or the per-micro-batch ready interval after the step starts. These two signals determine whether the trainer should spend GPUs on a decoupled Ref-Actor pipeline or release them to Rollout through colocated execution. Although StreamRL attempts to address this by dynamically increasing the size of Rollout, it lacks real-time reaction due to high initialization overheads. Furthermore, this scaling approach is incompatible with the fixed-resource environments prevalent in large-scale production clusters. As a result, the cluster is left severely underutilized throughout the training lifecycle, with goodput lost to both data starvation and off-policy throttling.

TIDE RL Overview

To address these challenges, we propose TIDE RL, an elastic asynchronous RL system tailored for multi-turn agentic workloads. TIDE RL coordinates three decisions that existing systems handle separately: which tasks stay resident in Rollout, how ready trajectories are consumed by the Trainer, and how GPUs move between the two sides.

Design Rationale. The design follows the dataflow of agentic RL. CTB keeps Rollout productive without losing task context; RA2P consumes ready micro-batches without waiting for a full global batch or thrashing between Ref and Actor; ERS adjusts resources so the selected RA2P strategy matches the current readiness pattern. The three components are intentionally coupled by one feedback loop: CTB controls when trajectories become ready, RA2P exposes how expensive it is to consume the current ready stream, and ERS moves ranks to reduce the dominant wait revealed by that stream.

System Architecture and Workflow. Inspired by AgentRL, TIDE RL adopts a disaggregated architecture where environment maintenance and reward computation are offloaded to an external CPU cluster. This isolates GPU instances purely for rollout and training. A task's lifecycle begins in the background, where CTB manages the admission of new agentic tasks based on Actor KV cache availability and strictly bounds off-policy staleness. Once dispatched to a Rollout rank, the Actor and Env engage in multi-turn interactions. During this highly concurrent multi-task generation, CTB monitors the real-time token footprint, pausing or resuming tasks to achieve higher throughput and prevent memory exhaustion. Upon completing all interaction turns, the Env calculates group-relative rewards and pushes the group to a global rollout buffer. The Trainer pulls ready trajectories from the buffer and uses RA2P to update the policy with low or no model-swapping overhead. Once the training step completes, Trainer synchronizes the new weights to Rollout. The overarching ERS scheduler monitors the utilization and readiness state of the rollout buffer. If trajectory generation outpaces training, ERS dynamically commands a subset of Rollout ranks to safely offload their Actor weights and reconfigures these freed nodes into the trainer group; if training starves, it moves trainer-side ranks back to Rollout and lets RA2P run in the colocated mode. TIDE RL also considers evaluation tasks. For the evaluation of the i-th step's model, CTB organizes them concurrently with training data rollout in the (i+1)-th step. CTB aggregates the metrics just before the weight synchronization of the (i+1)-th step. This ensures metrics are reported accurately without interrupting the continuous trajectory flow.

Continuous Task Batching

CTB makes Rollout scheduling task-aware. Instead of treating every environment turn as an independent request, CTB tracks each task's growing context, admits new tasks only when a Rollout rank has enough KV cache headroom, pauses tasks before memory is exhausted, and resumes paused tasks with cache locality whenever possible.

Token-Aware Admission Control. In highly concurrent agentic RL, dispatching tasks by raw request counts creates severe imbalance because tasks consume very different amounts of KV cache. CTB instead treats each Data Parallel (DP) rank as a token budget pool and admits tasks according to measured context size. To estimate a task's initial footprint, CTB sends a lightweight probing task for each trajectory group and records the token usage of its first turn. During dispatch, CTB checks the current token utilization of each rank and admits a task only if the remaining headroom can accommodate the probed footprint. After admission, CTB updates the task's actual token usage at the start of every interaction turn, so later scheduling decisions reflect the growing context rather than the initial estimate. CTB also enforces three admission constraints: Environment concurrency (Active tasks cannot exceed the parallel capacity of the external environments), Policy staleness (CTB throttles dispatch when Rollout is likely to generate more data than Trainer can consume within the staleness bound), and Task diversity (CTB penalizes assigning too many identical task types to the same batch, spreading environment-side stragglers across ranks).

Semantic-Aware Pausing and Resuming. As tasks progress through multiple turns, their contexts may outgrow a rank's token budget. CTB handles this by pausing selected tasks before KV cache pressure causes uncontrolled eviction. The priority score combines RL progress with system cost: P(t)=ω1·Ieval +ω2·Gcompletion +ω3·Iactive +ω4·Lcontext. Here, Ieval marks evaluation tasks, Gcompletion is the completion ratio of the task's GRPO group (e.g., Nfinished/Ntotal), Iactive marks tasks currently executing, and Lcontext is the current token length. The priority hierarchy is: Task Type (Evaluation tasks require anchored weights; stalling them blocks global weight synchronization - Highest priority), Group Comp. (GRPO requires full groups to calculate advantages; fragmented groups stall gradient updates - High priority), Exec. State (Preempting actively running tasks destroys in-flight compute cycles and prior investments - Medium priority), and Context Len. (Evicting massive historical contexts causes catastrophic redundant prefill upon resumption - Low priority). When capacity is restored, CTB resumes paused tasks before admitting new tasks. It also enforces worker affinity based on group IDs, placing resumed tasks on ranks that are most likely to retain their shared prefixes. This preserves useful KV cache state and avoids unnecessary prefix recomputation.

Resource-Aware Ref-Actor Collaboration

The Trainer must process rollout data as soon as it becomes useful, but agentic workloads make data readiness highly uneven. At some steps, many micro-batches are already waiting in the global buffer; at others, the Trainer starts almost empty and receives new micro-batches slowly. TIDE RL uses RAS (startup ready backlog) and TPRM (later ready interval) in two layers. First, RA2P provides two Ref-Actor execution strategies: a decoupled streaming strategy for high RAS or short TPRM, and a colocated aggregation strategy for low RAS and long TPRM. Second, ERS schedules between these two RA2P strategies by moving ranks between Rollout and Trainer.

Decoupled Streaming Mode. When RAS is large or TPRM is short, the Trainer has enough ready work to keep dedicated Ref and Actor resources busy. RA2P therefore avoids model-swapping overhead by allocating the Reference (Ref) and Actor models onto distinct GPUs. Since Ref only performs forward passes, RA2P groups one Ref rank with multiple Actor ranks to form a streaming micro-batch pipeline. To maximize concurrency and avoid blocking the Actor, we restructure the computation graph. In standard implementations, the loss computation is included in the Actor's forward pass, requiring Actor to wait for the Ref to output the reference log probabilities (ref log probs) before the forward pass. RA2P breaks this dependency by deferring the Actor's loss computation (including the KL divergence penalty) to its backward pass. The Ref model sequentially computes ref log probs, while the Actor performs its forward passes independently. The Actor only consumes the ref log probs when initiating the backward passes. This realignment ensures overlap between reference log prob computation and policy updates with much lower synchronization bubbles.

Optimized Colocated Mode. When RAS is small and TPRM is long, a dedicated Ref rank would spend much of the step waiting for data. RA2P then colocates the Ref and Actor models on the same GPUs, releasing the extra GPUs to Rollout where they can shorten TPRM for later steps. To keep colocation efficient, RA2P introduces two optimizations. First, we implement ready-batch aggregation. Instead of alternating between models for every single micro-batch (e.g., R1→A1→R2→A2), RA2P monitors the global rollout buffer. If multiple micro-batches are ready, it aggregates their execution. The system loads the Ref model once to process all ready batches, and then swaps to the Actor to perform updates. This amortizes the I/O cost of model alternation. Second, RA2P leverages zero-copy transmission. Because the Ref and Actor reside on the same physical node, the input sequences and the generated ref log probs are passed directly via GPU shared memory, bypassing network serialization and redundant memory allocations.

Analytical Mode Selection. The two strategies target different bottlenecks. Decoupled mode pays a pipeline-fill cost but removes model swaps, so it is best when high RAS amortizes the fill cost or short TPRM keeps the pipeline continuously fed. Colocated mode keeps fewer Trainer GPUs active and uses ready-batch aggregation to amortize occasional swaps, so it is best when low RAS and long TPRM would otherwise leave a decoupled Ref rank idle. RA2P compares the decoupled latency cost (Cdec) and colocated latency cost (Ccol) under these two signals.

Micro-Batch Dispatching. To maximize throughput and prevent stragglers, RA2P jointly optimizes how a micro-batch is formed and allocated across DP ranks. RA2P employs the same micro-batching strategy for both execution modes. Early micro-batches are accumulated from the rollout buffer until they contain sufficient tokens to fully saturate the compute capacity of all DP ranks. Conversely, the last micro-batch of a global batch is deliberately kept as small as possible. This intentionally minimizes the duration of the final forward and backward passes, thereby reducing the pipeline flush bubble at the end of the training step. RA2P further applies distinct dispatching strategies to handle the heterogeneous micro-batch size: In Colocated Mode, GPU execution is purely sequential. The primary goal is to balance the total compute load across all training ranks. RA2P utilizes a zig-zag (longest-processing-time-first) allocation strategy, sorting sequences by length and distributing them iteratively in a zig-zag pattern. This ensures that the sum of sequence lengths assigned to each rank remains nearly identical. In Decoupled Mode, the overall completion time is bottlenecked by the execution time of the final micro-batch. Building upon our small-final-batch sizing strategy, RA2P sorts and allocates the sequences such that the last sequence fed into the pipeline is the absolute smallest. This minimizes the tail latency, allowing the last active rank to finish its backward pass at the earliest possible time.

ERS Scheduling of RA2P Strategies

ERS closes the loop between RA2P's mode choice and the producer-consumer imbalance. It reallocates GPUs on the fly so that Rollout can generate data fast enough and Trainer can consume ready micro-batches with the right RA2P strategy.

Readiness-Aware Plan Generation. To accommodate volatile generation rates, TIDE RL adapts the elastic batching strategy, where the Trainer consumes any accumulated batch size falling within an interval [Bmin, Bideal]. Here, Bmin is the minimum batch size required for an effective training step, while Bideal represents the full utilization of GPU HBM and computing resources. ERS generates scaling plans immediately after the Trainer finishes an Actor update, before the new weights are broadcast (sync params). Despite CTB and RA2P, static resource allocation cannot perfectly align the generation and consumption throughputs. The mismatch appears as TWT when the Trainer starves for ready trajectories, and as RWT when Rollout is throttled by stale weights or an overfilled buffer. ERS monitors the global buffer to estimate the RAS and TPRM signals introduced above, and converts them into two complementary actions: scaling Rollout up to reduce TWT, and scaling Trainer up to reduce RWT. Directly using instantaneous arrivals yields a noisy signal skewed by trajectory variance even within one step. To avoid reactive thrashing, the ERS coordinator estimates these trends through two queue metrics: the generation and consumption deficits.

When Rollout resources are insufficient, generation slows. The next training step begins with a small RAS and observes a long TPRM, so the Trainer is starved both at startup and during the step, increasing TWT. ERS detects this via the generation deficit metric (the global batch volume <Bideal). It then selects colocated RA2P and reallocates the GPUs previously used by decoupled Ref ranks to Rollout. This gives generation more capacity, directly reducing TWT, while the Trainer avoids wasting dedicated Ref GPUs on an empty stream. Conversely, when Rollout is over-provisioned, ready data accumulates. The next training step starts with a large RAS, and continued fast generation shortens TPRM. If the Trainer cannot drain this data fast enough, Rollout eventually waits for buffer clearance and fresh weights, increasing RWT. ERS detects this consumption deficit through the Head-of-Line (HoL) latency of the oldest micro-batch in the global buffer, together with the current ready volume. If the HoL latency exceeds a threshold (τ×Tstep, where Tstep is the average step duration) or RAS exceeds the decoupled-mode break-even point, ERS selects decoupled RA2P and reclaims a Rollout rank to provision a dedicated Ref rank. The Trainer can then drain the backlog without repeated model swaps, directly reducing RWT.

Beyond standard training, TIDE RL handles boundary conditions via preemptive generation allocation. During cold starts or evaluation phases, generation dominates. Rather than waiting for RAS to remain low and TPRM to become long, the coordinator preemptively maximizes the Rollout group size, allocating maximum hardware to bootstrap the buffer quickly. TIDE RL adopts adaptive data retention to prevent starvation cycles following a Rollout scale-down. If a scaled-down Rollout cannot sustain the Trainer's consumption, the Trainer adaptively reduces its fetch size toward Bmin. By pacing consumption and retaining a micro-batch reservoir, ERS extends the step duration, granting Rollout time to accumulate a larger batch for the next iteration.

Seamless Scaling Operations. Traditional elastic frameworks suspend training to migrate model weights across nodes. TIDE RL eliminates this latency by exploiting the mathematical properties of on-policy RL. Iterative model weight updates render historical KV caches mathematically incompatible with the new policy. Thus, caches across Rollout ranks must be flushed at the post-update boundary. Piggybacking on this invalidation cycle, task migration becomes cache-free. TIDE RL bypasses physical KV cache transfers. When ERS closes a Rollout rank, the CTB scheduler automatically reassigns the active tasks on it to the remaining active ranks. Upon resumption, these tasks perform a fresh prefill using the synchronized new version of weights on their new Rollout ranks. To ensure dynamic resource reallocation does not block the training loop, TIDE RL overlaps PCIe weight-swapping operations with the sync params broadcast. When scaling down a Reference rank (to reassign it to Rollout), ERS offloads the reference model to the CPU while waiting for the Actor backward pass to complete. The reactivated Rollout rank then joins the sync params broadcast to load the Actor. Conversely, when scaling down a Rollout rank, it discards the stale Rollout model and skips the weight download. Instead, it utilizes the synchronization time window to load the Reference model concurrently with the other ranks' broadcast. Through this operational alignment, TIDE RL adjusts its functional GPU distribution without adding latency to the critical path.

Evaluation

We implement TIDE RL with ∼16000 lines of Python code, supporting mainstream training (e.g., Megatron, PyTorch FSDP) and serving (vLLM, SGLang) frameworks. We evaluate TIDE RL across text-only and multi-modal tasks, model sizes, and readiness regimes. Key findings are:

  • TIDE RL achieves over 5.6× RL training throughput on text-only tasks and reduces training time by 51.1% to reach similar task performance.

  • For multi-modal tasks, TIDE RL improves RL throughput by over 33% and reduces training time by 62.2% for similar task performance.

  • CTB improves KV cache hit rate by 1.58× and generation throughput by 1.15× by mitigating C1.

  • RA2P reduces per-step training time by up to 44.3% by selecting the proper ref-actor execution mode for different readiness patterns (C2).

  • ERS uses readiness signals to reallocate GPUs between rollout and training, reducing total waiting time by up to 77.6% (C3).

Methodology. We evaluate TIDE RL on a physically disaggregated cluster that completely separates the RL training backend from the environment simulation frontend. The training testbed consists of four compute nodes. Each node is equipped with 8× NVIDIA H100 GPUs interconnected via NVLink, a 64-core Intel Xeon CPU, and 1.5 TB of RAM memory. To support high-throughput distributed checkpointing and rapid parameter synchronization, a shared JuiceFS-backed NFS is deployed across the training cluster. We select a mixture of standard text-based and complex multimodal tasks. For text-based workloads, we employ a hybrid task suite comprising WebShop and AlfWorld. For multi-modal workloads, we utilize OSWorld and ScienceBoard. They require the agent to process high-resolution screenshots and interact with graphical user interfaces (GUIs), heavily stressing the prefill and context-carriage capacities of Rollout. We utilize a dedicated bare-metal Kubernetes cluster provisioned with 1024 CPU cores to run environments. We evaluate the system using a wide spectrum of state-of-the-art open-weights models. For text-only tasks, we employ the Qwen-2.5 series (7B and 14B). For multimodal tasks, we utilize advanced multi-modal language models, Qwen3-VL (4B) and Qwen-3.5 (9B). Qwen-3.5 also employs linear attention and multi-token prediction.

We compare TIDE RL with three RL training frameworks: VeRL (a highly optimized synchronous RL framework that enforces strict phase barriers), AReaL (a foundational asynchronous RL framework that physically decouples the Rollout and Trainer workers), and StreamRL (an advanced asynchronous framework featuring stream generation support). All frameworks use the same GPU budget, task stream, model checkpoints, staleness bound, global batch configuration, and rollout/training backends whenever the framework supports them. For fixed-partition asynchronous baselines, we sweep the rollout GPU ratio over the same candidate set and report the best-performing setting for each workload. We use training throughput as the primary system metric: the number of generated tokens that are eventually consumed by the Trainer per second. This excludes stale or discarded rollout tokens and therefore measures useful progress for on-policy RL.

End-to-End Performance on Text-Only Models. We run TIDE RL and the baselines for 100 training steps. Fixed-partition asynchronous systems use the best static rollout ratio from our sweep; on this workload, that ratio is 0.25. Compared with the synchronous VeRL baseline, TIDE RL achieves a 5.6× speedup. VeRL is bottlenecked by the slowest trajectories in each global batch: a wrong action can require extra environment turns and substantially extend rollout time. Its colocated synchronous execution also repeatedly swaps Rollout and Trainer states at every step, further reducing useful training throughput. AReaL improves over VeRL by decoupling generation from training, but its best fixed partition is still a single compromise across different readiness regimes. When evaluation or long-tail tasks make readiness sparse, Trainer stalls; when simple tasks make readiness dense, Rollout stalls behind the Trainer. TIDE RL reacts to these regimes through ERS and the two RA2P modes, achieving a 1.8× throughput improvement over AReaL. StreamRL reduces global-batch waiting by streaming micro-batches, but on text workloads each DP rank can receive more than 16 micro-batches per step. This bursty stream triggers frequent Ref-Actor model alternation and dominates the saved stall time; the 14B run does not complete within six hours. TIDE RL keeps the streaming benefit while avoiding model thrashing with RA2P, yielding over 7× speedup in this setting.

We further exhibit the Best-of-N (BoN) reward and the pass rate after the 100-th training step. The strictly synchronized VeRL has all the training data perfectly on policy with zero staleness. As expected, it exhibits the best training performance. Among the asynchronous RL frameworks, TIDE RL achieves the best task performance and remains close to VeRL, with a BoN reward deficit of 0.01 and a pass-rate deficit of 0.5%. It reaches this performance using only 48.9% of VeRL's wall-clock time. Compared with StreamRL and AReaL, TIDE RL generates and consumes useful trajectories faster, so more of its training data is consumed within the zero-staleness window. The other asynchronous baselines more often fall back to one-step-stale data, which explains why TIDE RL is closer to the synchronous learning curve while retaining much higher throughput.

End-to-End Performance on Multi-Modal Models. We run TIDE RL and the baselines on multi-modal tasks and evaluate both throughput and task performance. Compared with text-only workloads, multi-modal trajectories have longer environment interactions, larger observations, and more variable rollout/training balance. VeRL suffers more severely in this setting because long GUI interactions and high-variance observations amplify global-batch tail latency. All asynchronous frameworks improve over VeRL because sparse micro-batch arrivals make overlap more valuable. This regime also reduces StreamRL's model-swapping pressure compared with text tasks. Even after tuning AReaL's fixed ratio to provide sufficient Trainer capacity, TIDE RL achieves 6.02× throughput over VeRL and more than 1.33× over the asynchronous baselines. The gain is smaller than in text-only tasks because larger Trainer ranks reduce the number of ranks that ERS can migrate, but readiness-aware scheduling still avoids the worst fixed-partition stalls. After 40 training steps, TIDE RL consistently outperforms the three asynchronous baselines in BoN reward and pass rate. However, due to the high variance in training workloads, ERS generates deliberately conservative scaling plans. Because Trainer scale-up operations are bounded to one rank per step, the Trainer eventually emerges as the long-term pipeline bottleneck, causing the system to naturally gravitate toward a one-step off-policy staleness. While this staleness introduces a slight performance degradation compared to the fully synchronous VeRL, TIDE RL completes the 40 steps in only 37.8% of the wall-clock time, justifying the trade-off between algorithmic equivalence and training throughput.

Improvement Breakdown. To quantify the contribution of each TIDE RL component, we present an ablation study on OSWorld with Qwen-3-VL 4B. We use StreamRL as the baseline because it already streams rollout data to the Trainer, then add RA2P, ERS, and CTB cumulatively. StreamRL's main bottleneck is frequent model swapping across many micro-batches. RA2P removes PCIe transfers and GPU context-switching overhead from this path, improving Trainer throughput by 10.4%. ERS further adjusts GPU allocation: after the first step, it scales Rollout down and Trainer up, adding 35.7% improvement on top of RA2P by reducing both RWT and TWT. CTB improves the Rollout side by limiting harmful concurrency and reducing prefix recomputation, especially in the first step where many evaluation tasks run together. This yields a 54.8% first-step throughput improvement and a 4.8% overall improvement.

Extended Evaluation and Micro-Benchmarks. We isolate how each component addresses the bottlenecks. CTB improves KV cache hit rate and rollout throughput. We perform a step of 1,312 WebShop tasks on one H100. CTB improves cache hit rate by 1.58× and generation throughput by 1.15×, reducing step duration by 15.6%. The vanilla scheduler's frequent preemption drives the cache hit rate as low as 0.9%, while CTB keeps it above 6.0%, close to the ideal upper limit. RA2P reduces stall and thrashing. We show the overhead for processing eight ready micro-batches on four H100 GPUs with data parallelism only. Each micro-batch has 8,192 tokens. This setup isolates the high-RAS regime where the Trainer begins with enough work to expose model alternation overhead. We compare RA2P's colocated mode (CM) and decoupled mode (DM) against vanilla streaming with (VSO) and without (VS) model offloading. PCIe transfers and GPU context switching introduce significant overhead in the vanilla designs, while RA2P reduces overhead by up to 44.3%. ERS relieves both Rollout and Trainer waiting. We report the sum of RWT and TWT in the first 100 WebShop steps on one 8×NVIDIA H100 node, compared with the fixed rollout ratios. Among fixed settings, a rollout ratio of 0.25 gives the best throughput, but it still suffers high RWT in normal steps and high TWT during intermediate evaluation. These phases correspond to alternating large-RAS/short-TPRM and small-RAS/long-TPRM regimes. ERS moves GPUs between Rollout and Trainer and selects the matching RA2P mode, reducing total TWT and RWT by 68.6–77.6%.

Discussion

Incompatibility of Suffix Decoding. While suffix decoding (SD) effectively accelerates single-turn inference, our profiling shows that it is detrimental to highly variable multi-turn agentic workloads under our model scale. Enabling SD increases total rollout time regardless of CTB status. Complex environmental interactions yield low acceptance rates (∼40%) and short acceptance lengths (∼2.8 tokens), so draft-verification overhead outweighs speculative gains. Consequently, SD is excluded from our implementation.

Fault Tolerance and Recovery. In large-scale agentic RL, node failures are inevitable. TIDE RL provides robust fault tolerance without requiring global pipeline restarts. If a Rollout rank fails, the CTB scheduler immediately quarantines the node by halting new task dispatches, while the Trainer seamlessly excludes it from the asynchronous sync params broadcast. Furthermore, TIDE RL leverages asynchronous distributed checkpointing to periodically persist policy weights, ensuring rapid state recovery and minimal interruption to the global training loop.

Algorithmic Effectiveness. In TIDE RL, tokens within a single multi-turn trajectory might be generated by slightly different policy versions as weights update mid-rollout. Consistent with findings in asynchronous RL literature (e.g., APRIL and AReaL), this does not degrade algorithmic convergence. The KL penalty natively regularizes the policy against such version drifts. As demonstrated in §6, TIDE RL achieves highly competitive reward growth and learning efficiency compared to strict synchronous baselines.

Limitations and Future Work. While TIDE RL's algorithm design tolerates policy version drift during asynchronous training, evaluation tasks demand an anchored, static model version to ensure metric consistency. Currently, enforcing this strict version synchronization for evaluation can temporarily disrupt the asynchronous pipeline momentum. Future work could address this by integrating a relay weight synchronization mechanism (e.g., Laminar). This would allow the system to maintain decoupled, frozen model snapshots specifically for evaluation tasks without blocking the trainer, provided that the multi-version footprint in host memory can be carefully optimized to be kept with the reference model and large image files in trajectories simultaneously.

Conclusion

We present TIDE RL, an asynchronous system for multi-turn agentic RL. TIDE RL keeps Rollout efficient with task-level CTB, consumes ready trajectories with the two RA2P Ref-Actor strategies, and uses ERS to move GPUs between Rollout and Trainer as bottlenecks shift. On real testbeds, it improves training throughput by over 1.8× on text-only tasks and over 33% on multi-modal tasks compared with existing asynchronous systems.

Improvements for AI systems

Based on the paper, here are the specific improvements you can make to AI systems:

Improvement: Replace request-level continuous batching with task-level admission, preemption, and resumption that tracks token footprints per rank in real time.

What the improved system can do:

  • Maintain KV cache for multi-turn tasks across idle periods instead of evicting them

  • Prioritize tasks by semantic importance (evaluation boundaries > near-complete GRPO groups > active trajectories > long contexts)

  • Prevent catastrophic cache misses that cause redundant prefill computation

  • Achieve 1.58× higher KV cache hit rate and 1.15× higher generation throughput

Improvement: Dynamically select between two execution modes based on data readiness signals (RAS and TPRM):

  • Decoupled mode: Run Ref and Actor on separate GPUs with deferred loss computation in backward pass

  • Colocated mode: Use ready-batch aggregation and zero-copy shared-memory transfer

Improvement: Migrate GPU ranks between rollout and training on the fly using the same readiness signals, piggybacking weight migration onto natural cache-flush boundaries.

Improvement: Jointly optimize micro-batch formation and allocation across DP ranks with:

  • Small final batches to minimize pipeline flush bubbles

  • Zig-zag allocation for colocated mode (balance compute load)

  • Smallest-last sequencing for decoupled mode (minimize tail latency)

Improvement: Dispatch tasks based on measured token usage rather than raw request counts, with constraints on environment concurrency, policy staleness, and task diversity.

Improvement: Resume paused tasks on ranks most likely to retain shared prefixes, with worker affinity based on group IDs.

Improvement: When Rollout is scaled down, Trainer adaptively reduces fetch size toward Bmin and retains a micro-batch reservoir.

Improvement: During cold starts or evaluation phases, maximize Rollout group size preemptively rather than waiting for readiness signals.

Improvement: Overlap PCIe weight-swapping with sync params broadcast, discarding stale weights and loading new models during synchronization windows.

Improvement: Quarantine failed nodes, exclude them from broadcasts, and use asynchronous distributed checkpointing.

Overall system capability: The improved AI system can achieve 5.6× higher RL training throughput on text-only tasks and 33%+ on multi-modal tasks compared to existing systems, while reaching similar task performance with 51.1% less training time.

Abstract

Reinforcement learning (RL) for large language models is moving toward multi-turn agentic workloads, where rollout tasks repeatedly pause for external environments, resume with growing contexts, and finish at highly variable times. In this setting, RL training goodput, measured by training throughput, matters more than raw GPU occupancy: GPU waiting and repeated prefill recomputation are pure overhead. We present TideRL, a readiness-aware elastic RL system with Continuous Task Batching, Resource-Aware Ref-Actor Pipelining, and Elastic Resource Scaling. CTB preserves useful rollout state, RA 2 P selects between decoupled streaming and colocated aggregation from the ready backlog and arrival interval, and ERS moves ranks between rollout and training using the same readiness signals. Across text-only and multi-modal agentic workloads, TideRL improves RL training goodput by up to 5.6 times over synchronous baselines and over 33% over asynchronous baselines, while reaching similar task performance. It also improves KV cache hit rate by 1.58 times, reduces per-step training time by up to 44.3%, and cuts total waiting time by up to 77.6%.

Sources

Related papers