RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning
cs.CL, cs.AI
Submitted: 2026-09-17
Updated: 2026-09-27
Code: https://github.com/ZJU-REAL/SDAR
License: http://creativecommons.org/licenses/by/4.0/
The gist: Multi-turn agents trained with reinforcement learning (RL) receive a single scalar reward per trajectory, which motivates self on-policy distillation (OPD) to supply dense token-level supervision
Terminology
Abstract
Multi-turn agents trained with reinforcement learning (RL) receive a single scalar reward per trajectory, which motivates self on-policy distillation (OPD) to supply dense token-level supervision from a self-teacher with privileged task skills, letting a skill-free student internalize them. This recipe, however, is undermined by two findings in agentic tasks: privileged information alone does not always make a teacher reliable, and the benefit of teacher supervision is stage-dependent. We therefore propose RetireOPD (Self-Retiring On-Policy Distillation), which first optimizes a decoupled, skill-conditioned teacher with environment rewards and then trains a skill-free student jointly with RL and OPD. Rather than following a predefined distillation schedule, RetireOPD adopts Adaptive Retirement: the student drops the teacher on its own once their discrepancy stops shrinking and it reaches a target fraction of the teacher's success rate, after which training proceeds with RL alone. Across Qwen2.5 models from 1.5B to 7B, RetireOPD improves ALFWorld success rate over RL baseline by 14.1% to 18.8% and WebShop accuracy by 11.8% to 19.0%, and surpasses its own skill-conditioned teacher in every setting.
Sources
- HDPO: Hybrid Distillation Policy Optimization via Privileged Self-Distillation
- SAF-OPD: Stable Advantage Fusion for On-Policy Distillation
- Reinforcement Learning via Self-Distillation
- Hindsight Compacts but Does Not Repair: Rethinking On-Policy Self-Distillation in Reasoning Models
- Sequential Beats Joint: On the Interplay between On-Policy Distillation and RLVR
- Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
- UI-S1: Advancing GUI Automation via Semi-online Reinforcement Learning
- Self-Distilled Agentic Reinforcement Learning
- SKILL0: In-Context Agentic Reinforcement Learning for Skill Internalization
- Proximal Policy Optimization Algorithms
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- ALFWorld: Aligning Text and Embodied Environments for Interactive Learning
- OpenAI GPT-5 System Card
- ATOD: Annealed Turn-Aware On-Policy Distillation for Multi-Turn Agentic Tasks
- Qwen3.5-Omni Technical Report
- Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents
- TCOD: Exploring Temporal Curriculum in On-Policy Distillation for Multi-turn Autonomous Agents
- SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning
- MiMo-V2-Flash Technical Report
- DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering