Trajectory Learnability for Offline On-Policy Distillation with Imperfect Teachers
cs.LG, cs.AI
Submitted: 2026-09-16
Updated: 2026-09-16
Comments: 14 pages, 3 figures
License: http://creativecommons.org/licenses/by/4.0/
The gist: Offline on-policy distillation gains efficiency by collecting student trajectories and teacher supervision once and reusing them throughout optimization.
Terminology
Abstract
Offline on-policy distillation gains efficiency by collecting student trajectories and teacher supervision once and reusing them throughout optimization. The same reuse makes imperfect supervision persistent. Since even strong teachers can fail, we ask what remains learnable from imperfect teacher supervision? Teacher failure is only a coarse problem-level signal and does not imply that all supervision along the associated student trajectory is unhelpful. A natural alternative is to estimate teacher recoverability along the trajectory, but repeated continuations largely erase the efficiency advantage of offline distillation. We instead use teacher-successful problems to define a cheap reference for what the student can learn. We train on teacher-successful problems and measure how the likelihood of each observed token in trajectories from teacher-failed problems changes. We use these signed likelihood changes as an operational learnability signal: larger increases indicate behavior more strongly promoted by successful-only learning. We aggregate this signal into trajectory-level weights for the original distillation loss. Unlike continuation-based estimates, our learnability requires no additional generation and can be computed once from stored trajectories and model checkpoints. Across mathematical reasoning and code generation, our method improves an offline OPD baseline by up to 2.7 percentage points and matches or outperforms online OPD variants on multiple benchmarks. Despite the additional successful-only distillation stage, it uses 2 GPUs and about 22 GPU hours, compared with 3 GPUs and 36--48 GPU hours for representative online OPD methods.
Sources
- Training Verifiers to Solve Math Word Problems
- Keypoint-based Progressive Chain-of-Thought Distillation for LLMs
- When Teacher Guidance Misleads: Reward-Aligned On-Policy Distillation
- Distilling the Knowledge in a Neural Network
- Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
- Filter, Then Reweight: Rethinking Optimization Granularity in On-Policy Distillation
- ReNIO: Reweighting Negative Trajectory Importance for LLM On-Policy Distillation
- When Are Teacher Tokens Reliable? Position-Weighted On-Policy Self-Distillation for Reasoning
- Orca: Progressive Learning from Complex Explanation Traces of GPT-4
- Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning
- Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation
- Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation
- Scaling Relationship on Learning Mathematical Reasoning with Large Language Models
- Verify Before You Distill: Prompt-Level Teacher Gating for On-Policy Distillation
- ReOrder-OPD:Reliability-Aware Prompt Ordering for On-Policy Distillation
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks