Efficient Reinforcement Learning for Long-Horizon Tool-Use Agentic Tasks
Zelei Cheng, Amritansh Mishra, Sambit Sahu, William Campbell
Capital One
cs.LG, cs.AI
Submitted: 2026-08-11
Updated: 2026-08-12
Comments: Published at the COLM 2026 Workshop on Efficient Reasoning
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 75/100
The gist: S INK F LEX-RL is a modular training system for reinforcement learning (RL) in dual-control tool-use environments, designed to address the memory challenges of long-horizon agentic tasks.
Terminology
Summary
S INK F LEX-RL is a modular training system for reinforcement learning (RL) in dual-control tool-use environments, designed to address the memory challenges of long-horizon agentic tasks. The system integrates four key components: a Gymnasium-compatible environment wrapper, a VERL-style rollout dataflow, group-relative policy optimization (GRPO) without a separate value model, and a sink-aware FlexAttention path that preserves model-specific sink scaling under causal and sliding-window masks.
The paper addresses the systems side of memory-feasible agentic RL for a large open-weight mixture-of-experts (MoE) transformer. The setting stresses the training system through delayed verifiable rewards, multi-turn tool-heavy trajectories, and the need to coordinate environment execution, rollout generation, reward checking, and policy optimization. The environment is formalized as E = (g, πSOP, Atool, s0, u), where g is the user goal, πSOP denotes domain-policy constraints, Atool is the agent tool set, s0 is the initial shared state, and u is the user simulator. A trajectory is τ = (ot, at, rt, dt, it) tT=1, and the trainer consumes a trajectory-level reward Ri produced by the benchmark’s programmatic checker.
The training pipeline uses a common reset/step interface for each benchmark domain, with rollout workers owning the agent loop and formatting observations, sampling from the current policy, parsing responses into actions, and appending transitions to a trajectory buffer. The policy update uses GRPO, which computes a group-normalized trajectory advantage Âi = (Ri − µ(R1:G))/(σ(R1:G) + εA), where Ri is the programmatically computed trajectory reward. The loss minimized is LGRPO(θ) = −(1/∑i Ti)∑i∑t min(ρi,t(θ)Âi, clip(ρi,t(θ), 1−εc, 1+εc)Âi) + βDKL(πθ‖πref), where Ti is the optimized token length of rollout i, εc is the clipping radius, and the KL term regularizes the policy toward the reference model. The baseline design does not train a separate critic or value network.
The sink-aware FlexAttention path addresses the memory wall of quadratic attention score structures. The paper introduces a zero-value-sink equivalence: for a learned sink logit sη and value vector vsink = 0, the attention output with the explicit sink is algebraically equivalent to scaling the standard attention output by αsink = σ(l − sη), where l = log ∑i exp(q·ki). The implementation uses FlexAttention with a mask function Mb,h,q,k = 1[k ≤ q] ∧ 1[q − k ≤ w ∨ k < p], where w is the local-window size and p is the number of always-visible prefix positions. The attention call returns both the output and an auxiliary log-sum-exp statistic, and the sink path applies the model-specific scaling function to the output. The implementation composes FlexAttention with the sink-scaling operation under AOTAutograd and torch.compile, allowing the compiler to generate forward and backward code without materializing an O(n2) Jacobian. Memory-oriented optimizations include compilation and fusion of pointwise operations and mask broadcasting across batch and head dimensions.
The experiments report two complementary measurements. In the preliminary τ2-Bench retail training run, validation reward (mean@1) rises from 0.25 early in training to 0.44 later in the observed training window, while training-score and trajectory-reward proxies rise from 0.18 to 0.40 and 0.39, respectively. These values are visually estimated from dashboard screenshots and represent two portions of the same run rather than an untrained baseline and a final converged model. The paper notes that without multiple seeds, an optimizer baseline, or exported scalar logs, these observations do not isolate the effect of GRPO, establish variance reduction, or support a statistical significance claim.
In the peak-memory study, the optimized sink-aware FlexAttention path reduces peak VRAM at every sequence length for which both paths complete. At 4096 tokens, the optimized path reduces peak VRAM from 28.06 GB to 22.52 GB, a reduction of 5.54 GB or 19.7%. At 8192 tokens, it completes the measured configuration with a peak allocation of 25.53 GB, whereas the eager reference path encounters an out-of-memory error. The paper emphasizes that because the experiment measures only peak VRAM, it does not establish improvements in throughput, latency, total training time, or accelerator utilization.
The discussion highlights that the system is particularly well suited to agentic tasks with programmatically verifiable outcomes, as instantiated by τ2-Bench. The preliminary retail-domain experiment provides encouraging evidence for the effectiveness of the proposed policy-learning pipeline, with displayed reward-associated traces exhibiting a clear upward trend. Programmatic rewards provide an efficient and scalable source of supervision, and GRPO is well matched to settings where rollout groups contain diverse behavioral outcomes. The memory evaluation demonstrates feasibility at sequence lengths up to 8192 tokens under the tested configuration, establishing an important systems proof of concept.
The conclusion states that the results support the narrower claim that environment, RL-dataflow, and attention-kernel integration can improve the memory feasibility of long-horizon agent training. They do not yet establish algorithmic superiority, broad generalization, exact implementation equivalence, or end-to-end computational speedup. The ethics statement notes that the system trains agents that can call tools and update environment state, and such agents should be evaluated for policy compliance, user deception, unsafe tool use, privacy leakage, and simulator overfitting before deployment.
Improvements for AI systems
Improvements to AI systems based on this paper:
- Memory-efficient long-context RL training for tool-use agents
-
Integrate the sink-aware FlexAttention path (with zero-value-sink equivalence and sliding-window masks) into any transformer-based RL trainer. This allows training on trajectories up to 8192 tokens without OOM, reducing peak VRAM by 20% at 4096 tokens and enabling longer rollouts that were previously infeasible.
-
The improved system can train MoE models on multi-turn, tool-heavy agentic tasks (e.g., web navigation, API calls, database queries) with delayed rewards, without needing a separate critic network.
- Scalable GRPO without value models for programmatic-reward tasks
-
Adopt the group-relative policy optimization (GRPO) loss with trajectory-level rewards from programmatic checkers, eliminating the need for learned value functions. This reduces memory and compute overhead while still providing stable policy updates.
-
The improved system can handle sparse, verifiable rewards (e.g., task completion, correctness of tool outputs) in environments with diverse rollout groups, enabling faster iteration on benchmarks like τ2-Bench.
- Modular environment-rollout-reward pipeline for dual-control tool use
-
Use the Gymnasium-compatible wrapper and VERL-style rollout dataflow to decouple environment simulation, policy sampling, and reward checking. This allows parallel rollout workers to own the agent loop, format observations, parse actions, and buffer trajectories efficiently.
-
The improved system can be rapidly adapted to new tool-use benchmarks by swapping only the environment wrapper and reward checker, without rewriting the RL training core.
- Compiler-optimized attention with auxiliary log-sum-exp for sink scaling
-
Implement the FlexAttention composition with AOTAutograd and torch.compile to generate fused forward/backward kernels, avoiding O(n2) Jacobian materialization. The auxiliary log-sum-exp statistic enables model-specific sink scaling without extra memory.
-
The improved system can train with causal and sliding-window masks at longer sequence lengths, preserving model-specific attention biases (e.g., learned sink tokens) while maintaining numerical equivalence to explicit sink implementations.
- Memory-feasible training for long-horizon agentic RL
-
Combine the above components to train agents on trajectories with many steps (e.g., 100+ tool calls) and delayed rewards, where full-sequence attention would otherwise be prohibitive.
-
The improved system can now perform RL fine-tuning of large open-weight MoE models (e.g., 70B+ parameters) on agentic tasks using a single GPU with 32GB VRAM, up to 8192-token sequences, enabling research and deployment in resource-constrained settings.
- Systematic evaluation for verifiable agentic tasks
-
Leverage the paper’s emphasis on programmatic rewards and trajectory-level advantages to build RL systems that are more sample-efficient and less prone to reward hacking, since rewards are based on objective task completion rather than learned heuristics.
-
The improved system can be evaluated on benchmarks with clear success criteria, allowing for reproducible comparisons across seeds and baselines (as the paper notes, future work should include multiple seeds and optimizer baselines).
Abstract
Long-horizon tool-using agents must reason over user goals, domain policies, tool calls, simulator state, and delayed verifiable rewards. Reinforcement learning (RL) is a natural fit for this setting, but multi-turn on-policy rollouts create long contexts, while model-specific attention layers may require custom masks and learned sink normalization. We present SINKFLEX-RL, a modular training system for RL in dual-control tool-use environments. The system combines a Gymnasium-compatible environment wrapper, a VERL-style rollout dataflow, group-relative policy optimization without a separate value model, and a sink-aware FlexAttention path designed to preserve model-specific sink scaling under causal and sliding-window masks. In a preliminary Tau2Bench retail run, validation reward (mean@1) rises from 0.25 early in training to 0.44 later in the observed training window, while training-score and trajectory-reward proxies also trend upward. In a fixed-configuration memory benchmark, the optimized attention path reduces peak VRAM from 28.06GB to 22.52GB at 4096 tokens, a 19.7% reduction, and runs the measured 8192-token configuration using 25.53 GB where the eager baseline runs out of memory. These results illustrate the value of integrating environment interfaces, RL dataflow, and attention-kernel design for memory-feasible long-horizon agent training.
Sources
- $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment
- Proximal Policy Optimization Algorithms
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- HybridFlow: A Flexible and Efficient RLHF Framework
- $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks