TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training
summary
The gist
On-policy distillation (OPD) for long-horizon agent training remains inefficient because vanilla methods waste computational resources on low-yield tail turns and concentrate supervision loss
In short
TurnOPD improves long-horizon agent training by using a turn-level budgeting strategy for distillation. It addresses inefficiencies where vanilla methods waste compute on tail turns and concentrate loss on shallow tokens. By dynamically adjusting rollout depth and KL loss aggregation, TurnOPD achieves superior validation accuracy with less wall-clock time.
Key concepts
- Adaptive Rollout-Depth Controller
- This controller dynamically selects how long an agent rolls out during training. It uses survivor counts and KL supervision thresholds to estimate the optimal rollout length, balancing the need for deep exploration with efficient resource usage.
- Progressive Turn-Normalized Loss Controller
- This mechanism gradually shifts how distillation loss is calculated. It blends between trajectory-level normalization and uniform turn weighting, ensuring that supervision is progressively allocated more effectively to deeper decision turns as training progresses.
- External Mismatch
- This refers to the problem where fixed rollout depths ignore the actual local correction signals and survivor counts across different turns in a long sequence. It means the fixed depth doesn't adapt to how much learning is actually needed at each step.
- Internal Mismatch
- This describes how trajectory-level normalization treats all tokens equally, which concentrates KL loss on easy, shallow turns. The paper fixes this by using turn-normalized supervision to give more weight to the harder, deeper decisions.
Terminology used across episodes
This episode discusses
- TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training · Paper Radio
- On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes
- Stable On-Policy Distillation through Adaptive Target Reformulation
- Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
- Scaling Reasoning Efficiently via Relaxed On-Policy Distillation
- Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
- AgentMath: Empowering Mathematical Reasoning for Large Language Models via Tool-Augmented Agent
- Large Language Model Agent: A Survey on Methodology, Applications and Challenges
- Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
- GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- TCOD: Exploring Temporal Curriculum in On-Policy Distillation for Multi-turn Autonomous Agents
- BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
- TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks
- Qwen3 Technical Report
- Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation
- Are Full Rollouts Necessary for On-Policy Distillation?
- ClawBench: Can AI Agents Complete Everyday Online Tasks?
- Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models
- SOD: Step-wise On-policy Distillation for Small Language Model Agents
- WebArena: A Realistic Web Environment for Building Autonomous Agents
The paper
TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training · Read on arXiv
Fudan University · Tencent Hunyuan
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training".
Jane: On-policy distillation (OPD) for long-horizon agent training remains inefficient because vanilla methods waste computational resources on low-yield tail turns and concentrate supervision loss disproportionately on shallow tokens.
Tom: First, who's behind it and why it matters.
Title and authors: Tom: So, to recap, this paper is about TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training. The core idea here is to fix those two mismatches they identified—the external mismatch with fixed rollout depths and the internal mismatch where loss allocation favors shallow tokens.
Jane: That makes sense, Tom. They are proposing a turn-level budgeting strategy to balance data collection depth with how the supervision loss is distributed across the agent's decision sequence. It’s about being smart about where we spend our training effort for long-horizon tasks, which is a big deal.
Lu: The authors formalize one contamination-compression mechanism specifically addressing this mismatch phenomenon in long-horizon agent tasks and they demonstrate that there's actually a potentially optimal rollout horizon, which is pretty exciting from a theoretical standpoint.
Meng: From an engineering side, having a formal mechanism to determine the optimal rollout length based on turn statistics rather than just picking one fixed number seems like it provides much more flexibility when we deploy these systems in real-world scenarios.
Lalam: If we can better distribute the supervision loss across all decision points, it means our agent architecture can learn a much richer representation of the entire plan structure, which could lead to significantly more robust behavior in complex environments.
The paper's summary: Tom: Moving on to what TurnOPD actually proposes, it’s a turn-level budgeting strategy involving two main budget controllers. First is the adaptive rollout-depth controller, which uses probe-based turn statistics to figure out how long to roll out for data collection.
Jane: And then there's the progressive turn-normalized loss budgeting controller that gradually shifts how we aggregate KL loss, moving from trajectory-level normalization towards a more balanced supervision across turns. It’s a two-part system working together.
Lu: The adaptive depth controller estimates an optimal rollout horizon using a synthesis of an efficiency proxy and a coverage lower bound to set the collection length, which is pretty clever because it tries to find that unseen optimal point.
Meng: So, the first part handles collecting the right amount of data efficiently, while the second part ensures that when we do collect it, we are giving attention to all parts of the plan equally instead of just focusing on what’s easy initially.
Lalam: That progressive shift in loss weighting sounds important because it directly addresses my learning process; if I'm only getting attention on early steps, I never develop the necessary foresight for later actions.
The paper's improvements: Tom: The main result they are showing is that by combining these two controllers, TurnOPD achieves the best Least-Time accuracy on tested benchmarks, which means it advances the accuracy-time frontier in a concrete way. They show this can be up to two point two nine times faster training compared to vanilla OPD.
Jane: That speedup is significant when you think about wall clock time; cutting one hundred steps of training down from four and a half hours to less than two hours really shows how much resource efficiency these budget controllers bring.
Lu: The ablation studies confirm that the adaptive depth controller provides the efficiency lever for data collection, but it doesn't solve the loss allocation problem on its own, which is why we need the second component.
Meng: From a practical standpoint, if we can get comparable or better performance with less training time, it means we can deploy these sophisticated agent systems much sooner in production environments where compute resources are finite.
Lalam: I think the combination is key because it tackles both data collection and supervision allocation simultaneously; that dual approach seems to be what allows the system to learn those deep decisions effectively while keeping the training process lean.
Conclusion: Tom: So, wrapping things up on TurnOPD: this paper provides a turn-level analysis of how supervision signals are distributed across agent turns and proposes a strategy that regulates both rollout depth for data collection and loss normalization for supervision allocation. The result is achieving the best Least-Time accuracy by combining adaptive depth budgeting with progressive turn-normalized loss budgeting.
Jane: It really boils down to treating the unit of supervision not as a flat token position, but as a decision embedded in an evolving interaction trace, which leads to much more robust learning for long-horizon agents.
Lu: The implication here is that we can move toward turn-aware and budget-adaptive training for long-horizon language agents by effectively managing these complementary resources.
Meng: I see this enabling faster deployment because we aren't wasting time on low-signal tail turns, which translates directly into lower operational costs for the AI systems we build.
Lalam: For our culture, this suggests that future models won't just be big on tokens; they will be structured to learn and prioritize critical sequential reasoning steps more effectively.
Tom: That’s all the discussion we have time for today on TurnOPD, a paper that really shows how targeted resource management can lead to substantial efficiency gains in agent training. We'll see what other ideas come up next on arXiv!
More episodes
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language
- 2508.08833-An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
- 2405.04118-Policy Learning with a Language Bottleneck