TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training

arXiv:2607.05804 · cs.AI, cs.CL · Submitted 2026-07-07 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training".

Jane: On-policy distillation (OPD) for long-horizon agent training remains inefficient because vanilla methods waste computational resources on low-yield tail turns and concentrate supervision loss disproportionately on shallow tokens.

Tom: First, who's behind it and why it matters.

Title and authors: Tom: So, to recap, this paper is about TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training. The core idea here is to fix those two mismatches they identified—the external mismatch with fixed rollout depths and the internal mismatch where loss allocation favors shallow tokens.

Jane: That makes sense, Tom. They are proposing a turn-level budgeting strategy to balance data collection depth with how the supervision loss is distributed across the agent's decision sequence. It’s about being smart about where we spend our training effort for long-horizon tasks, which is a big deal.

Lu: The authors formalize one contamination-compression mechanism specifically addressing this mismatch phenomenon in long-horizon agent tasks and they demonstrate that there's actually a potentially optimal rollout horizon, which is pretty exciting from a theoretical standpoint.

Meng: From an engineering side, having a formal mechanism to determine the optimal rollout length based on turn statistics rather than just picking one fixed number seems like it provides much more flexibility when we deploy these systems in real-world scenarios.

Lalam: If we can better distribute the supervision loss across all decision points, it means our agent architecture can learn a much richer representation of the entire plan structure, which could lead to significantly more robust behavior in complex environments.

The paper's summary: Tom: Moving on to what TurnOPD actually proposes, it’s a turn-level budgeting strategy involving two main budget controllers. First is the adaptive rollout-depth controller, which uses probe-based turn statistics to figure out how long to roll out for data collection.

Jane: And then there's the progressive turn-normalized loss budgeting controller that gradually shifts how we aggregate KL loss, moving from trajectory-level normalization towards a more balanced supervision across turns. It’s a two-part system working together.

Lu: The adaptive depth controller estimates an optimal rollout horizon using a synthesis of an efficiency proxy and a coverage lower bound to set the collection length, which is pretty clever because it tries to find that unseen optimal point.

Meng: So, the first part handles collecting the right amount of data efficiently, while the second part ensures that when we do collect it, we are giving attention to all parts of the plan equally instead of just focusing on what’s easy initially.

Lalam: That progressive shift in loss weighting sounds important because it directly addresses my learning process; if I'm only getting attention on early steps, I never develop the necessary foresight for later actions.

The paper's improvements: Tom: The main result they are showing is that by combining these two controllers, TurnOPD achieves the best Least-Time accuracy on tested benchmarks, which means it advances the accuracy-time frontier in a concrete way. They show this can be up to two point two nine times faster training compared to vanilla OPD.

Jane: That speedup is significant when you think about wall clock time; cutting one hundred steps of training down from four and a half hours to less than two hours really shows how much resource efficiency these budget controllers bring.

Lu: The ablation studies confirm that the adaptive depth controller provides the efficiency lever for data collection, but it doesn't solve the loss allocation problem on its own, which is why we need the second component.

Meng: From a practical standpoint, if we can get comparable or better performance with less training time, it means we can deploy these sophisticated agent systems much sooner in production environments where compute resources are finite.

Lalam: I think the combination is key because it tackles both data collection and supervision allocation simultaneously; that dual approach seems to be what allows the system to learn those deep decisions effectively while keeping the training process lean.

Conclusion: Tom: So, wrapping things up on TurnOPD: this paper provides a turn-level analysis of how supervision signals are distributed across agent turns and proposes a strategy that regulates both rollout depth for data collection and loss normalization for supervision allocation. The result is achieving the best Least-Time accuracy by combining adaptive depth budgeting with progressive turn-normalized loss budgeting.

Jane: It really boils down to treating the unit of supervision not as a flat token position, but as a decision embedded in an evolving interaction trace, which leads to much more robust learning for long-horizon agents.

Lu: The implication here is that we can move toward turn-aware and budget-adaptive training for long-horizon language agents by effectively managing these complementary resources.

Meng: I see this enabling faster deployment because we aren't wasting time on low-signal tail turns, which translates directly into lower operational costs for the AI systems we build.

Lalam: For our culture, this suggests that future models won't just be big on tokens; they will be structured to learn and prioritize critical sequential reasoning steps more effectively.

Tom: That’s all the discussion we have time for today on TurnOPD, a paper that really shows how targeted resource management can lead to substantial efficiency gains in agent training. We'll see what other ideas come up next on arXiv!

Fudan University · Tencent Hunyuan

cs.AI, cs.CL

Submitted: 2026-07-07

Updated: 2026-09-28

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 83/100

The gist: On-policy distillation (OPD) for long-horizon agent training remains inefficient because vanilla methods waste computational resources on low-yield tail turns and concentrate supervision loss

Key concepts

Adaptive Rollout-Depth Controller
This controller dynamically selects how long an agent rolls out during training. It uses survivor counts and KL supervision thresholds to estimate the optimal rollout length, balancing the need for deep exploration with efficient resource usage.
Progressive Turn-Normalized Loss Controller
This mechanism gradually shifts how distillation loss is calculated. It blends between trajectory-level normalization and uniform turn weighting, ensuring that supervision is progressively allocated more effectively to deeper decision turns as training progresses.
External Mismatch
This refers to the problem where fixed rollout depths ignore the actual local correction signals and survivor counts across different turns in a long sequence. It means the fixed depth doesn't adapt to how much learning is actually needed at each step.
Internal Mismatch
This describes how trajectory-level normalization treats all tokens equally, which concentrates KL loss on easy, shallow turns. The paper fixes this by using turn-normalized supervision to give more weight to the harder, deeper decisions.

Terminology

Summary

On-policy distillation (OPD) for long-horizon agent training remains inefficient because vanilla methods waste computational resources on low-yield tail turns and concentrate supervision loss disproportionately on shallow tokens. TurnOPD addresses this by proposing a turn-level budgeting strategy that dynamically adjusts both rollout depth and KL loss aggregation to create a more balanced supervision signal.

The gist: TurnOPD proposes a turn-level budgeting strategy for efficient on-policy distillation of long-horizon agents, consisting of adaptive rollout-depth budgeting and progressive turn-normalized loss budgeting, which achieves superior validation accuracy under equal wall-clock training budgets and advances the accuracy–time frontier beyond vanilla OPD.

Diagnosis of Inefficiencies

The paper identifies two key inefficiencies in vanilla agent OPD when applied to long-horizon agent tasks: (1) full-horizon rollouts often waste wall-clock resources on tail turns that provide weak and noisy KL supervision, and (2) trajectory-level KL objectives concentrate most of the loss on shallow tokens, leaving deeper decision turns under-trained once initial behaviors are aligned. These issues stem from two mismatches: an External mismatch where fixed rollout depth ignores varying local correction signals and survivor counts across turns, and an Internal mismatch where trajectory-level normalization gives uniform token weights, concentrating the KL signal on easy, shallow turns.

TurnOPD Budget Controllers

TurnOPD consists of two budget controllers designed to regulate complementary resources:

  1. Adaptive rollout-depth controller: This controller dynamically selects rollout length via periodic probes, guided by survivor-weighted KL and coverage thresholds. It estimates the optimal (but unobservable) rollout horizon, denoted as H⋆, by synthesizing an efficiency-centric proxy (Heff) and a coverage-based lower bound (Hcov), where Hctrl = max(Heff, Hcov).

  2. Progressive turn-normalized loss controller: This controller gradually transitions KL loss aggregation from trajectory-level to turn-balanced supervision. It uses a linear blend between trajectory-level normalization and uniform turn weighting: q blend t = (1 − α) q traj t + α q turn t, where the blend coefficient α is tied to normalized training progress, enabling better learning of deep decisions.

Experimental Results and Performance

Experiments on ALFWorld, WebShop, and Multi-Hop Search show that TurnOPD achieves superior validation accuracy under equal wall-clock training budgets. For example, on ALFWorld-1.7B, it increases Same-Step Avg@4 from 83.0 to 86.3 and cuts 100-step wall time from 4.42h to 1.93h, representing a speedup of up to 2.29x compared to vanilla OPD. TurnOPD is demonstrated as the best Least-Time accuracy method across tested benchmarks, advancing the accuracy–time frontier by combining the efficiency lever (adaptive depth) and the optimization lever (loss allocation).

Ablation Studies and Controller Alignment

Ablation studies confirm that both budget controllers contribute uniquely to performance. Adaptive depth alone saves compute but does not fix the loss allocation problem. The linear KL blend improves accuracy by reallocating supervision toward deeper turns, showing that the rollout-depth budget controller provides the efficiency lever, the progressive loss-budget controller provides the optimization lever, and their combination gives the best accuracy–time tradeoff. Furthermore, analysis of KL normalization schemes shows that a linear blend offers a smoother transition than a hard turn-level KL switch, yielding better Same-Step Avg@4 by balancing stability and targeted allocation.

Conclusion

The work concludes that the unit of supervision in long-horizon agents should not be treated as a flat token position, but as a turn-conditioned decision embedded in an evolving interaction trace. TurnOPD successfully regulates complementary resources—rollout depth for data collection and loss normalization for supervision allocation—to yield robust accuracy improvements with less computation. The proposed strategy provides a path toward turn-aware and budget-adaptive training of long-horizon language agents.

References

[1] Rishabh Agarwal, Nino Vieillard, Yongchao Zhou, Piotr Stanczyk, Sabela Ramos, Matthieu Geist, and Olivier Bachem. On-policy distillation of language models: Learning from self-generated mistakes. 2024.

[15] Kevin Lu and Thinking Machines Lab. On-policy distillation. Thinking Machines Lab: Connectionism, 2025.

[16] Haipeng Luo, Huawen Feng, Qingfeng Sun, Can Xu, Kai Zheng, Yufei Wang, Tao Yang, Han Hu, and Yansong Tang. Agentmath: Empowering mathematical reasoning for large language models via tool-augmented agent. arXiv preprint arXiv:2512.20745.

Improvements for AI systems

As a fastidious and diligent researcher, I have analyzed TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training. The core contribution is a turn-level budgeting strategy that addresses two fundamental allocation mismatches in vanilla on-policy distillation (OPD) for long-horizon agents: external mismatch (fixed rollout depth vs. variable turn utility) and internal mismatch (trajectory normalization concentrating loss on shallow tokens).

Here are the specific improvements and what the resulting AI system can achieve:


The proposed TurnOPD framework enhances AI systems by optimizing the training process for complex, multi-turn agentic tasks, leading to significant gains in both accuracy and computational efficiency.

  1. A turn-aware budget controller dynamically manages data collection depth based on real-time supervision signals rather than fixed rollout lengths.

  2. A progressive loss allocation strategy ensures that optimization efforts are distributed equitably across all critical decision points (turns) within the agent's interaction trace, regardless of token count.

The improved AI system can perform the following specific actions:

  1. An agent can execute long, complex plans (e.g., multi-step web navigation, multi-hop search queries) with a significantly higher success rate compared to vanilla OPD methods under the same wall-clock training budget.

  2. The agent will reach state of high accuracy faster by efficiently allocating training resources to the most informative turns—the deeper decision points—rather than wasting compute on low-signal or noisy tail turns.

  3. The system achieves superior Accuracy–Time Frontier, meaning it can achieve a target performance level (e.g., 90% accuracy) in substantially less wall-clock time (up to 2.29x faster training, as shown on ALFWorld).

  4. The system will exhibit better robustness across different task types (embodied planning, web navigation, and information retrieval) because the turn-level budgeting adapts its strategy based on task-specific signals (e.g., how teacher uncertainty or outcome separation changes over depth).

  5. The agent's policy refinement will be more targeted; deep decision turns—which are crucial for long-horizon reasoning—will receive a proportionally larger share of the optimization loss, leading to better learning of complex, late-stage planning behaviors.

Abstract

On-policy distillation (OPD) trains a student policy by matching a stronger teacher on the student's own trajectories, offering a promising framework for language agent training. However, its application to long-horizon agentic tasks remains insufficiently explored. We identify two key inefficiencies in vanilla agent OPD: (1) full-horizon rollouts often waste wall-clock resources on tail turns that provide weak and noisy KL supervision, and (2) trajectory-level KL objectives concentrate most of the loss on shallow tokens, leaving deeper decision turns under-trained once initial behaviors are aligned. To address these challenges, we propose TurnOPD, a turn-level budgeting strategy for efficient on-policy distillation of long-horizon agents. TurnOPD consists of two budget controllers: adaptive rollout-depth budgeting, which uses probe-based turn statistics to determine rollout length, and progressive turn-normalized loss budgeting, which gradually shifts KL weighting from token-level to turn-balanced supervision. Experiments on ALFWorld, WebShop, and Multi-Hop Search with task-specialized teacher models show that TurnOPD achieves superior validation accuracy under equal wall-clock training budgets and advances the accuracy--time frontier beyond vanilla OPD.

Sources

Related papers