TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training".
Jane: On-policy distillation (OPD) for long-horizon agent training remains inefficient because vanilla methods waste computational resources on low-yield tail turns and concentrate supervision loss disproportionately on shallow tokens.
Tom: First, who's behind it and why it matters.
Title and authors: Tom: So, to recap, this paper is about TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training. The core idea here is to fix those two mismatches they identified—the external mismatch with fixed rollout depths and the internal mismatch where loss allocation favors shallow tokens.
Jane: That makes sense, Tom. They are proposing a turn-level budgeting strategy to balance data collection depth with how the supervision loss is distributed across the agent's decision sequence. It’s about being smart about where we spend our training effort for long-horizon tasks, which is a big deal.
Lu: The authors formalize one contamination-compression mechanism specifically addressing this mismatch phenomenon in long-horizon agent tasks and they demonstrate that there's actually a potentially optimal rollout horizon, which is pretty exciting from a theoretical standpoint.
Meng: From an engineering side, having a formal mechanism to determine the optimal rollout length based on turn statistics rather than just picking one fixed number seems like it provides much more flexibility when we deploy these systems in real-world scenarios.
Lalam: If we can better distribute the supervision loss across all decision points, it means our agent architecture can learn a much richer representation of the entire plan structure, which could lead to significantly more robust behavior in complex environments.
The paper's summary: Tom: Moving on to what TurnOPD actually proposes, it’s a turn-level budgeting strategy involving two main budget controllers. First is the adaptive rollout-depth controller, which uses probe-based turn statistics to figure out how long to roll out for data collection.
Jane: And then there's the progressive turn-normalized loss budgeting controller that gradually shifts how we aggregate KL loss, moving from trajectory-level normalization towards a more balanced supervision across turns. It’s a two-part system working together.
Lu: The adaptive depth controller estimates an optimal rollout horizon using a synthesis of an efficiency proxy and a coverage lower bound to set the collection length, which is pretty clever because it tries to find that unseen optimal point.
Meng: So, the first part handles collecting the right amount of data efficiently, while the second part ensures that when we do collect it, we are giving attention to all parts of the plan equally instead of just focusing on what’s easy initially.
Lalam: That progressive shift in loss weighting sounds important because it directly addresses my learning process; if I'm only getting attention on early steps, I never develop the necessary foresight for later actions.
The paper's improvements: Tom: The main result they are showing is that by combining these two controllers, TurnOPD achieves the best Least-Time accuracy on tested benchmarks, which means it advances the accuracy-time frontier in a concrete way. They show this can be up to two point two nine times faster training compared to vanilla OPD.
Jane: That speedup is significant when you think about wall clock time; cutting one hundred steps of training down from four and a half hours to less than two hours really shows how much resource efficiency these budget controllers bring.
Lu: The ablation studies confirm that the adaptive depth controller provides the efficiency lever for data collection, but it doesn't solve the loss allocation problem on its own, which is why we need the second component.
Meng: From a practical standpoint, if we can get comparable or better performance with less training time, it means we can deploy these sophisticated agent systems much sooner in production environments where compute resources are finite.
Lalam: I think the combination is key because it tackles both data collection and supervision allocation simultaneously; that dual approach seems to be what allows the system to learn those deep decisions effectively while keeping the training process lean.
Conclusion: Tom: So, wrapping things up on TurnOPD: this paper provides a turn-level analysis of how supervision signals are distributed across agent turns and proposes a strategy that regulates both rollout depth for data collection and loss normalization for supervision allocation. The result is achieving the best Least-Time accuracy by combining adaptive depth budgeting with progressive turn-normalized loss budgeting.
Jane: It really boils down to treating the unit of supervision not as a flat token position, but as a decision embedded in an evolving interaction trace, which leads to much more robust learning for long-horizon agents.
Lu: The implication here is that we can move toward turn-aware and budget-adaptive training for long-horizon language agents by effectively managing these complementary resources.
Meng: I see this enabling faster deployment because we aren't wasting time on low-signal tail turns, which translates directly into lower operational costs for the AI systems we build.
Lalam: For our culture, this suggests that future models won't just be big on tokens; they will be structured to learn and prioritize critical sequential reasoning steps more effectively.
Tom: That’s all the discussion we have time for today on TurnOPD, a paper that really shows how targeted resource management can lead to substantial efficiency gains in agent training. We'll see what other ideas come up next on arXiv!
Fudan University · Tencent Hunyuan
cs.AI, cs.CL
Submitted: 2026-07-07
Updated: 2026-09-28
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 83/100
The gist: On-policy distillation (OPD) for long-horizon agent training remains inefficient because vanilla methods waste computational resources on low-yield tail turns and concentrate supervision loss
Key concepts
- Adaptive Rollout-Depth Controller
- This controller dynamically selects how long an agent rolls out during training. It uses survivor counts and KL supervision thresholds to estimate the optimal rollout length, balancing the need for deep exploration with efficient resource usage.
- Progressive Turn-Normalized Loss Controller
- This mechanism gradually shifts how distillation loss is calculated. It blends between trajectory-level normalization and uniform turn weighting, ensuring that supervision is progressively allocated more effectively to deeper decision turns as training progresses.
- External Mismatch
- This refers to the problem where fixed rollout depths ignore the actual local correction signals and survivor counts across different turns in a long sequence. It means the fixed depth doesn't adapt to how much learning is actually needed at each step.
- Internal Mismatch
- This describes how trajectory-level normalization treats all tokens equally, which concentrates KL loss on easy, shallow turns. The paper fixes this by using turn-normalized supervision to give more weight to the harder, deeper decisions.
Terminology
Summary
On-policy distillation (OPD) for long-horizon agent training remains inefficient because vanilla methods waste computational resources on low-yield tail turns and concentrate supervision loss disproportionately on shallow tokens. TurnOPD addresses this by proposing a turn-level budgeting strategy that dynamically adjusts both rollout depth and KL loss aggregation to create a more balanced supervision signal.
The gist: TurnOPD proposes a turn-level budgeting strategy for efficient on-policy distillation of long-horizon agents, consisting of adaptive rollout-depth budgeting and progressive turn-normalized loss budgeting, which achieves superior validation accuracy under equal wall-clock training budgets and advances the accuracy–time frontier beyond vanilla OPD.
Diagnosis of Inefficiencies
The paper identifies two key inefficiencies in vanilla agent OPD when applied to long-horizon agent tasks: (1) full-horizon rollouts often waste wall-clock resources on tail turns that provide weak and noisy KL supervision, and (2) trajectory-level KL objectives concentrate most of the loss on shallow tokens, leaving deeper decision turns under-trained once initial behaviors are aligned. These issues stem from two mismatches: an External mismatch
where fixed rollout depth ignores varying local correction signals and survivor counts across turns, and an Internal mismatch
where trajectory-level normalization gives uniform token weights, concentrating the KL signal on easy, shallow turns.
TurnOPD Budget Controllers
TurnOPD consists of two budget controllers designed to regulate complementary resources:
-
Adaptive rollout-depth controller: This controller dynamically selects rollout length via periodic probes, guided by
survivor-weighted KL and coverage thresholds.
It estimates the optimal (but unobservable) rollout horizon, denoted as H⋆, by synthesizing anefficiency-centric proxy
(Heff) and acoverage-based lower bound
(Hcov), where Hctrl = max(Heff, Hcov). -
Progressive turn-normalized loss controller: This controller gradually transitions KL loss aggregation from trajectory-level to turn-balanced supervision. It uses a linear blend between trajectory-level normalization and uniform turn weighting: q blend t = (1 − α) q traj t + α q turn t, where the blend coefficient α is tied to normalized training progress, enabling
better learning of deep decisions.
Experimental Results and Performance
Experiments on ALFWorld, WebShop, and Multi-Hop Search show that TurnOPD achieves superior validation accuracy under equal wall-clock training budgets. For example, on ALFWorld-1.7B, it increases Same-Step Avg@4 from 83.0 to 86.3 and cuts 100-step wall time from 4.42h to 1.93h, representing a speedup of up to 2.29x compared to vanilla OPD. TurnOPD is demonstrated as the best Least-Time accuracy
method across tested benchmarks, advancing the accuracy–time frontier by combining the efficiency lever (adaptive depth) and the optimization lever (loss allocation).
Ablation Studies and Controller Alignment
Ablation studies confirm that both budget controllers contribute uniquely to performance. Adaptive depth alone saves compute but does not fix the loss allocation problem. The linear KL blend improves accuracy by reallocating supervision toward deeper turns, showing that the rollout-depth budget controller provides the efficiency lever, the progressive loss-budget controller provides the optimization lever, and their combination gives the best accuracy–time tradeoff.
Furthermore, analysis of KL normalization schemes shows that a linear blend
offers a smoother transition than a hard turn-level KL switch,
yielding better Same-Step Avg@4 by balancing stability and targeted allocation.
Conclusion
The work concludes that the unit of supervision in long-horizon agents should not be treated as a flat token position, but as a turn-conditioned decision embedded in an evolving interaction trace.
TurnOPD successfully regulates complementary resources—rollout depth for data collection and loss normalization for supervision allocation—to yield robust accuracy improvements with less computation. The proposed strategy provides a path toward turn-aware and budget-adaptive training of long-horizon language agents.
References
[1] Rishabh Agarwal, Nino Vieillard, Yongchao Zhou, Piotr Stanczyk, Sabela Ramos, Matthieu Geist, and Olivier Bachem. On-policy distillation of language models: Learning from self-generated mistakes. 2024.
[15] Kevin Lu and Thinking Machines Lab. On-policy distillation. Thinking Machines Lab: Connectionism, 2025.
[16] Haipeng Luo, Huawen Feng, Qingfeng Sun, Can Xu, Kai Zheng, Yufei Wang, Tao Yang, Han Hu, and Yansong Tang. Agentmath: Empowering mathematical reasoning for large language models via tool-augmented agent. arXiv preprint arXiv:2512.20745.
Improvements for AI systems
As a fastidious and diligent researcher, I have analyzed TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training.
The core contribution is a turn-level budgeting strategy that addresses two fundamental allocation mismatches in vanilla on-policy distillation (OPD) for long-horizon agents: external mismatch (fixed rollout depth vs. variable turn utility) and internal mismatch (trajectory normalization concentrating loss on shallow tokens).
Here are the specific improvements and what the resulting AI system can achieve:
The proposed TurnOPD framework enhances AI systems by optimizing the training process for complex, multi-turn agentic tasks, leading to significant gains in both accuracy and computational efficiency.
-
A turn-aware budget controller dynamically manages data collection depth based on real-time supervision signals rather than fixed rollout lengths.
-
A progressive loss allocation strategy ensures that optimization efforts are distributed equitably across all critical decision points (turns) within the agent's interaction trace, regardless of token count.
The improved AI system can perform the following specific actions:
-
An agent can execute long, complex plans (e.g., multi-step web navigation, multi-hop search queries) with a significantly higher success rate compared to vanilla OPD methods under the same wall-clock training budget.
-
The agent will reach state of high accuracy faster by efficiently allocating training resources to the most informative turns—the deeper decision points—rather than wasting compute on low-signal or noisy tail turns.
-
The system achieves superior
Accuracy–Time Frontier,
meaning it can achieve a target performance level (e.g., 90% accuracy) in substantially less wall-clock time (up to 2.29x faster training, as shown on ALFWorld). -
The system will exhibit better robustness across different task types (embodied planning, web navigation, and information retrieval) because the turn-level budgeting adapts its strategy based on task-specific signals (e.g., how teacher uncertainty or outcome separation changes over depth).
-
The agent's policy refinement will be more targeted; deep decision turns—which are crucial for long-horizon reasoning—will receive a proportionally larger share of the optimization loss, leading to better learning of complex, late-stage planning behaviors.
Abstract
On-policy distillation (OPD) trains a student policy by matching a stronger teacher on the student's own trajectories, offering a promising framework for language agent training. However, its application to long-horizon agentic tasks remains insufficiently explored. We identify two key inefficiencies in vanilla agent OPD: (1) full-horizon rollouts often waste wall-clock resources on tail turns that provide weak and noisy KL supervision, and (2) trajectory-level KL objectives concentrate most of the loss on shallow tokens, leaving deeper decision turns under-trained once initial behaviors are aligned. To address these challenges, we propose TurnOPD, a turn-level budgeting strategy for efficient on-policy distillation of long-horizon agents. TurnOPD consists of two budget controllers: adaptive rollout-depth budgeting, which uses probe-based turn statistics to determine rollout length, and progressive turn-normalized loss budgeting, which gradually shifts KL weighting from token-level to turn-balanced supervision. Experiments on ALFWorld, WebShop, and Multi-Hop Search with task-specialized teacher models show that TurnOPD achieves superior validation accuracy under equal wall-clock training budgets and advances the accuracy--time frontier beyond vanilla OPD.
Sources
- On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes
- Stable On-Policy Distillation through Adaptive Target Reformulation
- Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
- Scaling Reasoning Efficiently via Relaxed On-Policy Distillation
- Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
- AgentMath: Empowering Mathematical Reasoning for Large Language Models via Tool-Augmented Agent
- Large Language Model Agent: A Survey on Methodology, Applications and Challenges
- Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
- GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- TCOD: Exploring Temporal Curriculum in On-Policy Distillation for Multi-turn Autonomous Agents
- BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
- TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks
- Qwen3 Technical Report
- Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation
- Are Full Rollouts Necessary for On-Policy Distillation?
- ClawBench: Can AI Agents Complete Everyday Online Tasks?
- Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models
- SOD: Step-wise On-policy Distillation for Small Language Model Agents
- WebArena: A Realistic Web Environment for Building Autonomous Agents
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection