Latent On-Policy Self-Distillation

arXiv:2608.13040 · cs.LG, cs.CL · Submitted 2026-08-13 · Read on arXiv

Guibin Zhang, Jiayang Lyu, Ran Sun, Xinlei Yu, Haoyu Zhao, Qibing Ren, Shuicheng Yan

National University of Singapore · Beijing University of Posts and Telecommunications · Shanghai Jiao Tong University

cs.LG, cs.CL

Submitted: 2026-08-13

Updated: 2026-08-14

Code: https://github.com/bingreeky/LOPD

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 95/100

The gist: The paper addresses the challenge of enabling agents to learn from experience and internalize it into their policy for self-evolving AI.

Terminology

Summary

The paper addresses the challenge of enabling agents to learn from experience and internalize it into their policy for self-evolving AI. On-policy self-distillation (OPSD) provides a pathway by using a privileged self-teacher to offer dense supervision on the student's own trajectories. However, existing OPSD methods still rely heavily on designer-specified privileged artifacts (e.g., answers, feedback, skills, or trajectories), limiting the end-to-end learnability and scalability required for continual self-improvement.

The central research question posed is: "Can privileged context itself be learned end-to-end from experience, so that the self-teacher, rather than a designer-specified rule, determines what experiential knowledge to retain and how to encode it for dense on-policy supervision?"

LOPD makes the teacher's privileged context itself learnable end-to-end from experience. Technically, LOPD "retrieves relevant experiences and composes them into continuous latent tokens that condition a self-teacher, while the student generates trajectories from the task and interaction history and receives dense token-level supervision at every visited prefix."

The method introduces a privileged-margin objective to stabilize and regulate the learning of latent context. The key innovation is that the teacher is conditioned not on a pre-defined privileged artifact, but on learnable latent context instantiated by a composer that transforms retrieved experiences into compact continuous tokens.

The composer consists of an encoder (frozen backbone with trainable LoRA adapter) and a QFormer-style compressor with cross-attention. Retrieved experiences are encoded into hidden states, then compressed into K=32 latent tokens per experience (with J=3 retrieved experiences, totaling 96 latent tokens). The teacher receives these latent tokens as ordinary context positions, while the student conditions only on the task and interaction history.

The training pipeline follows standard OPSD: "the student first rolls out trajectories from the task and interaction history, and the teacher re-evaluates every visited prefix to provide dense token-level distributions, against which the student is optimized through reverse-KL distillation."

To prevent the teacher from collapsing toward the student, LOPD imposes a privileged-margin constraint that ties the teacher's token-level advantage to the verified outcome of the complete trajectory. The constraint requires the teacher to maintain a verifiable log-probability advantage over the student, with the dual variable β updated to enforce a minimum privilege level m=0.05. The paper notes: "If the composer degenerates to uninformative context, πT → πS, δt,n → 0, ∆ → 0 < m, and the dual penalty activates—structurally excluding the trivial solution."

LOPD obtains the best aggregate result in all ten backbone–benchmark comparisons. Key results include:

  • Tool use with Qwen3-4B: LOPD improves EnvScaler from 61.8 to 63.7, BFCL-v3 from 25.25 to 27.38, and ACEBench from 56.0 to 60.6 over the strongest competing method.

  • Tool use with Qwen3-8B: LOPD reaches 66.4/29.88/62.7 on EnvScaler, BFCL-v3, and ACEBench, compared with strongest baseline results of 60.2/29.00/58.0.

  • Code generation with Qwen3-4B: LOPD improves LiveCodeBench and EvalPlus aggregates to 48.78 and 81.36.

  • Code generation with Olmo3-7B: LOPD reaches 50.98 and 78.41, outperforming strongest alternatives by 2.69 and 0.55 points.

LOPD surpasses GRPO and Skill-SD with less than 30% of their rollout budget. In training dynamics, LOPD exceeds 0.61 mean reward after 320 generations and reaches 0.637 by generation 576, while GRPO and Skill-SD improve more gradually and finish at 0.611 and 0.588.

The paper provides direct evidence that making privileged context learnable is necessary for realizing these gains. Key findings:

  • Training with a frozen composer achieves 0.573 on EnvScaler.

  • Without the margin constraint (m=0), the student drops to 0.551.

  • With m≥0.02, the student surpasses the frozen-composer baseline, reaching 0.637 at m=0.05.

The distilled student inherits the interaction pattern induced by latent context: it uses more environment steps (17.04 vs. 11.12) but makes far fewer tool calls per step (1.11 vs. 3.50), has a 37.5% shorter first-step response, and reduces repeated calls from 8.89 to 5.25.

The paper demonstrates that no hand-crafted context is universally optimal—adding privileged information does not always improve upon the vanilla model. For example, SDPO falls from 22.88 to 15.75 on BFCL-v3 and OPSD likewise reduces the LiveCodeBench aggregate to 40.24. LOPD avoids these issues by learning a compact continuous representation directly from retrieved trajectories.

The paper positions LOPD as a step toward a more scalable and self-directed paradigm for agent evolution. The central claim is that "scalable self-evolution should not depend on a succession of increasingly elaborate, human-authored experience formats. Raw trajectories provide a minimal substrate, and richer repositories or retrievers may broaden what is available, but learning should decide what becomes useful guidance."

Improvements for AI systems

Improvements to AI systems:

  1. Self-evolving agents with learnable memory composition: Replace hand-crafted privileged context (e.g., predefined skills, feedback templates, or trajectory formats) with a trainable composer that retrieves relevant past experiences and compresses them into continuous latent tokens. The agent learns end-to-end which experiences matter and how to encode them for dense supervision, eliminating designer bias and enabling adaptation to novel tasks without manual prompt engineering.

  2. Dense token-level self-distillation with outcome-verified margins: Implement a teacher-student framework where the teacher conditions on latent context and provides token-level supervision at every visited prefix, while a privileged-margin constraint (dual variable β, minimum margin m=0.05) ties the teacher’s advantage to verified trajectory outcomes. This prevents teacher collapse toward the student and ensures the distilled policy internalizes only verifiable improvements, not noise.

  3. Sample-efficient policy optimization via latent context: Use LOPD’s training loop (student rollouts → teacher re-evaluation → reverse-KL distillation) to surpass methods like GRPO and Skill-SD with less than 30% of their rollout budget. The improved system reaches higher mean rewards faster (0.637 vs. 0.611 for GRPO by generation 576), reducing compute and data requirements for continual learning.

  4. Behavioral regularization through distilled interaction patterns: The distilled student inherits latent-context-induced behaviors—e.g., using more environment steps (17.04 vs. 11.12) but fewer tool calls per step (1.11 vs. 3.50), shorter first-step responses (37.5% reduction), and fewer repeated calls (5.25 vs. 8.89). This yields more deliberate, less redundant action sequences, improving reliability in tool-use and code-generation tasks.

  5. Robustness to suboptimal privileged information: The system avoids performance degradation seen when adding hand-crafted context (e.g., SDPO drops from 22.88 to 15.75 on BFCL-v3; OPSD reduces LiveCodeBench to 40.24). LOPD’s learned latent context ensures that even when retrieved experiences are imperfect, the composer filters and compresses them into useful guidance, maintaining or improving performance across all benchmarks.

What the improved AI system can do:

  • Continually self-improve from raw trajectory data alone, without human-authored skill libraries, feedback formats, or answer templates, by learning what to retain and how to encode it.

  • Achieve state-of-the-art results in tool use (e.g., EnvScaler 66.4, BFCL-v3 29.88, ACEBench 62.7 with Qwen3-8B) and code generation (e.g., LiveCodeBench 50.98, EvalPlus 78.41 with Olmo3-7B), outperforming strong baselines across ten backbone–benchmark pairs.

  • Learn faster with fewer rollouts, making it practical for real-world deployment where interaction budgets are limited.

  • Exhibit more efficient and interpretable behavior—fewer redundant tool calls, shorter responses, and more deliberate exploration—leading to lower latency and cost in production agent systems.

  • Generalize across diverse tasks without manual per-task context design, as the composer adapts the latent representation to the task and interaction history, avoiding the pitfalls of fixed privileged artifacts.

Abstract

Enabling agents to learn from experience and internalize it into their policy has become a central problem in self-evolving AI. On-policy self-distillation (OPSD) offers an effective pathway by using a privileged self-teacher to provide dense supervision on the student's own trajectories; however, existing methods still rely heavily on designer-specified privileged artifacts (e.g., answers, feedback, skills, or trajectories), limiting the end-to-end learnability and scalability required for continual self-improvement. In this work, we introduce Latent On-Policy Self-Distillation (LOPD), which, rather than proposing another hand-crafted OPSD variant with a newly prescribed form of privileged context, makes the teacher's privileged context itself learnable end-to-end from experience. Technically, LOPD retrieves relevant experiences and composes them into continuous latent tokens that condition a self-teacher, while the student generates trajectories from the task and interaction history and receives dense token-level supervision at every visited prefix. We further introduce a privileged-margin objective to stabilize and regulate the learning of latent context. Empirically, LOPD demonstrates (I) strong performance, outperforming RLVR and representative OPSD methods including OPSD, SDPO, and Skill-SD across both agentic tool use and code generation; and (II) high learning efficiency, surpassing GRPO and Skill-SD with less than 30% of their rollout budget. Ablation studies further provide direct evidence that making privileged context learnable is necessary for realizing these gains. Together, these results position LOPD as a step toward a more scalable and self-directed paradigm for agent evolution.

Sources

Related papers