Beyond Token-Local Imitation: Reward-Compatible Temporal Credit Assignment for On-Policy Distillation
cs.LG, cs.AI, cs.PL
Submitted: 2026-09-15
Updated: 2026-09-19
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: On-policy distillation (OPD) has emerged as an effective approach for large language model post-training, yet existing objectives face a trade-off between objective fidelity and optimization
Terminology
Abstract
On-policy distillation (OPD) has emerged as an effective approach for large language model post-training, yet existing objectives face a trade-off between objective fidelity and optimization stability. Token-level OPD provides stable but local supervision, whereas sequence-level OPD captures future credit at the cost of horizon-dependent variance. We establish a unified temporal-credit view of these formulations, showing that practical token-level OPD can be interpreted as a temporal approximation to the sequence-level reverse-KL gradient. Building on this connection, we propose γ OPD, which uses discounted temporal credit assignment to balance long-horizon supervision and optimization stability, while admitting a horizon-independent variance bound. We further develop a reward-compatible bounded mixing (RBM) mechanism for γ OPD that balances verifiable outcome feedback with the discounted OPD advantage to move beyond purely teacher-dependent optimization. Experiments on mathematical and code reasoning demonstrate consistent improvements over existing OPD methods across vanilla, size-mismatched, and multi-teacher distillation settings.
Sources
- Process Reinforcement through Implicit Rewards
- DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
- SAF-OPD: Stable Advantage Fusion for On-Policy Distillation
- Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes
- DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
- Asymmetric On-Policy Distillation: Bridging Exploitation and Imitation at the Token Level
- Entropy-Aware On-Policy Distillation of Language Models
- Kimi K3: Open Frontier Intelligence
- STAPO: Stabilizing Reinforcement Learning for LLMs by Silencing Rare Spurious Tokens
- KL for a KL: On-Policy Distillation with Control Variate Baseline
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Qwen3 Technical Report
- Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation
- Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy Distillation
- Are Full Rollouts Necessary for On-Policy Distillation?
- SCOPE: Signal-Calibrated On-Policy Distillation Enhancement with Dual-Path Adaptive Weighting
- Less is More: Early Stopping Rollout for On-Policy Distillation
- Reward-Weighted On-Policy Distillation with an Open Property-Equivalence Verifier for NL-to-SVA Generation
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks