Towards Understanding On-Policy Distillation through the Lens of Test-Time Scaling
Shanghai Jiao Tong University · Shanghai Innovation Institute · Ant Group · Hong Kong Baptist University · University of Texas at Austin
cs.LG, cs.CL
Submitted: 2026-08-12
Updated: 2026-09-28
Comments: 15 pages, 8 figures
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 95/100
Terminology
Summary
Summary
This paper investigates whether on-policy distillation (OPD) truly expands the reasoning capability boundary of student language models or primarily improves sampling efficiency within existing capabilities. The authors examine OPD through the lens of test-time scaling by varying the sampling budget K and evaluating performance with pass@K and avg@K across three student-teacher settings (Qwen3, Skywork, JustRL) and four reasoning benchmarks (AMC2023, AIME2024/2025/2026).
Key findings:
-
OPD improves sampling efficiency but does not expand the capability boundary. The paper states:
OPD-trained models perform better pass@K at small K values, but are surpassed by pre-OPD base models at large K, while consistently attaining higher avg@K across different sampling budgets.
At K=1024, OPD-trained models do not exceed their pre-OPD base models across all settings and benchmarks. -
OPD forgets more solvable problems than it newly learns. Using pass@1024 as the solvability criterion, the paper finds:
after OPD training, more previously solvable problems become unsolvable than previously unsolvable problems become solvable.
The retained fraction increases with K and becomes substantially larger than either learned or forgotten fractions. -
A trade-off emerges during OPD training. The paper observes:
as OPD training progresses, pass@K at small K values (e.g., K = 1 or 2) generally improves, while pass@K at large K (e.g., K = 512 or 1024) gradually decreases.
The degradation of the capability boundary is unstable and emerges early, around 80 training steps. -
Advanced OPD variants show similar behavior. EOPD, ExOPD, Direct-OPD, and pure forward-KL variants
achieve higher pass@K at small K and higher avg@K across sampling budgets
butat larger sampling budgets, their pass@K values are generally below or match those of the pre-OPD base model.
-
Off-policy distillation differs fundamentally. The paper states:
Off-policy distillation actually transfers teacher's capability to student, yet OPD not.
Off-policy distillation improves pass@K across both small and large sampling budgets, whereas OPD does not. -
Perplexity analysis confirms the interpretation. Reasoning trajectories generated by OPD-trained models are more favored by the teacher model than those from the base model, but "OPD primarily shifts probability mass toward reasoning trajectories that are already well supported by the pre-OPD base model, rather than introducing reasoning patterns from the teacher that are unfamiliar to the pre-OPD base model."
The paper concludes that OPD exhibits an illusory distillation
effect: "despite leveraging a stronger teacher model, OPD primarily improves the accessibility of reasoning capabilities already present in the student, rather than consistently transferring new capabilities beyond the student's original boundary."
Improvements for AI systems
Improvements to AI systems:
-
Add capability-boundary-aware training curricula. Modify the distillation training loop to periodically evaluate the student model at large sampling budgets (e.g., pass@512 or pass@1024) on a held-out reasoning benchmark. If the large-K performance drops below its pre-training baseline, halt or reduce OPD updates and switch to off-policy distillation or direct fine-tuning on teacher-generated solutions for previously unsolvable problems. This prevents the
illusory distillation
effect where the model becomes more efficient at generating known solutions but loses rare, hard-won reasoning paths. -
Implement a dual-objective loss with a
forgetting penalty.
During OPD, add a regularization term that penalizes the model when it decreases the log-probability of solutions that were solvable at high sampling budgets before training (e.g., using a reference log-probability from the pre-OPD model). This directly counteracts the observed trade-off where small-K performance improves at the cost of large-K capability degradation, preserving the model's original reasoning boundary while still gaining sampling efficiency. -
Introduce a sampling-budget-aware inference controller. After OPD training, deploy a meta-controller that dynamically selects between the OPD-trained model and the pre-OPD base model based on the available inference budget. For low-budget settings (K ≤ 16), use the OPD model for faster, more accurate answers; for high-budget settings (K ≥ 256), fall back to the base model to retain access to rare but correct solutions that OPD tends to forget. This hybrid approach maximizes both efficiency and ceiling performance.
-
Build a
solvability-tracking
diagnostic tool for model evaluation. Create a monitoring system that, during any distillation or fine-tuning run, tracks per-problem solvability at multiple K values (e.g., K=1, 16, 1024) and reports the learned/forgotten/retained fractions in real time. This allows practitioners to detect capability-boundary erosion early (the paper shows it emerges around 80 steps) and intervene before irreversible forgetting occurs, rather than relying on final aggregate metrics that mask the trade-off. -
Develop a teacher-consistency filter for distillation data. Before applying OPD, filter the teacher-generated reasoning trajectories to include only those that are not already well-supported by the student’s current policy (e.g., using a threshold on the student’s log-probability). The paper shows OPD mainly reinforces existing patterns; by explicitly selecting unfamiliar teacher trajectories, the student is forced to learn genuinely new reasoning structures, potentially converting OPD from illusory to true capability transfer.
What the improved AI system can do:
-
Maintain or expand its maximum reasoning capability (pass@1024) while still improving average performance and low-budget accuracy, avoiding the silent degradation of hard-problem solving.
-
Adaptively allocate inference compute — use fast, distilled models for easy/medium queries and switch to the original model for hard queries, ensuring no loss of ceiling performance.
-
Provide early warnings during training if the model is about to forget rare solutions, allowing automatic rollback or curriculum adjustment.
-
Learn genuinely new reasoning patterns from the teacher rather than merely re-weighting existing ones, leading to actual capability expansion in domains where the student was previously weak.
-
Achieve both efficiency and robustness — faster responses for typical tasks without sacrificing the ability to solve extremely difficult problems that require extensive sampling.
Sources
- Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
- Evaluating Large Language Models Trained on Code
- Training Verifiers to Solve Math Word Problems
- DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
- Weak-to-Strong Generalization via Direct On-Policy Distillation
- JustRL: Scaling a 1.5B LLM with a Simple RL Recipe
- Skywork Open Reasoner 1 Technical Report
- Distilling the Knowledge in a Neural Network
- Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
- The Flexibility Trap: Rethinking the Value of Arbitrary Order in Diffusion Language Models
- A Survey of On-Policy Distillation for Large Language Models
- Demystifying On-Policy Distillation: Roles, Pathologies, and Regulations
- MiMo-V2-Flash Technical Report
- Trust Region On-Policy Distillation
- A Survey on Knowledge Distillation of Large Language Models
- Qwen3 Technical Report
- Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation
- Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy Distillation
- DAPO: An Open-Source LLM Reinforcement Learning System at Scale
- The Many Faces of On-Policy Distillation: Pitfalls, Mechanisms, and Fixes
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks