Towards Understanding On-Policy Distillation through the Lens of Test-Time Scaling

arXiv:2608.11829 · cs.LG, cs.CL · Submitted 2026-08-12 · Read on arXiv

Shanghai Jiao Tong University · Shanghai Innovation Institute · Ant Group · Hong Kong Baptist University · University of Texas at Austin

cs.LG, cs.CL

Submitted: 2026-08-12

Updated: 2026-09-28

Comments: 15 pages, 8 figures

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 95/100

Terminology

Summary

Summary

This paper investigates whether on-policy distillation (OPD) truly expands the reasoning capability boundary of student language models or primarily improves sampling efficiency within existing capabilities. The authors examine OPD through the lens of test-time scaling by varying the sampling budget K and evaluating performance with pass@K and avg@K across three student-teacher settings (Qwen3, Skywork, JustRL) and four reasoning benchmarks (AMC2023, AIME2024/2025/2026).

Key findings:

  1. OPD improves sampling efficiency but does not expand the capability boundary. The paper states: OPD-trained models perform better pass@K at small K values, but are surpassed by pre-OPD base models at large K, while consistently attaining higher avg@K across different sampling budgets. At K=1024, OPD-trained models do not exceed their pre-OPD base models across all settings and benchmarks.

  2. OPD forgets more solvable problems than it newly learns. Using pass@1024 as the solvability criterion, the paper finds: after OPD training, more previously solvable problems become unsolvable than previously unsolvable problems become solvable. The retained fraction increases with K and becomes substantially larger than either learned or forgotten fractions.

  3. A trade-off emerges during OPD training. The paper observes: as OPD training progresses, pass@K at small K values (e.g., K = 1 or 2) generally improves, while pass@K at large K (e.g., K = 512 or 1024) gradually decreases. The degradation of the capability boundary is unstable and emerges early, around 80 training steps.

  4. Advanced OPD variants show similar behavior. EOPD, ExOPD, Direct-OPD, and pure forward-KL variants achieve higher pass@K at small K and higher avg@K across sampling budgets but at larger sampling budgets, their pass@K values are generally below or match those of the pre-OPD base model.

  5. Off-policy distillation differs fundamentally. The paper states: Off-policy distillation actually transfers teacher's capability to student, yet OPD not. Off-policy distillation improves pass@K across both small and large sampling budgets, whereas OPD does not.

  6. Perplexity analysis confirms the interpretation. Reasoning trajectories generated by OPD-trained models are more favored by the teacher model than those from the base model, but "OPD primarily shifts probability mass toward reasoning trajectories that are already well supported by the pre-OPD base model, rather than introducing reasoning patterns from the teacher that are unfamiliar to the pre-OPD base model."

The paper concludes that OPD exhibits an illusory distillation effect: "despite leveraging a stronger teacher model, OPD primarily improves the accessibility of reasoning capabilities already present in the student, rather than consistently transferring new capabilities beyond the student's original boundary."

Improvements for AI systems

Improvements to AI systems:

  1. Add capability-boundary-aware training curricula. Modify the distillation training loop to periodically evaluate the student model at large sampling budgets (e.g., pass@512 or pass@1024) on a held-out reasoning benchmark. If the large-K performance drops below its pre-training baseline, halt or reduce OPD updates and switch to off-policy distillation or direct fine-tuning on teacher-generated solutions for previously unsolvable problems. This prevents the illusory distillation effect where the model becomes more efficient at generating known solutions but loses rare, hard-won reasoning paths.

  2. Implement a dual-objective loss with a forgetting penalty. During OPD, add a regularization term that penalizes the model when it decreases the log-probability of solutions that were solvable at high sampling budgets before training (e.g., using a reference log-probability from the pre-OPD model). This directly counteracts the observed trade-off where small-K performance improves at the cost of large-K capability degradation, preserving the model's original reasoning boundary while still gaining sampling efficiency.

  3. Introduce a sampling-budget-aware inference controller. After OPD training, deploy a meta-controller that dynamically selects between the OPD-trained model and the pre-OPD base model based on the available inference budget. For low-budget settings (K ≤ 16), use the OPD model for faster, more accurate answers; for high-budget settings (K ≥ 256), fall back to the base model to retain access to rare but correct solutions that OPD tends to forget. This hybrid approach maximizes both efficiency and ceiling performance.

  4. Build a solvability-tracking diagnostic tool for model evaluation. Create a monitoring system that, during any distillation or fine-tuning run, tracks per-problem solvability at multiple K values (e.g., K=1, 16, 1024) and reports the learned/forgotten/retained fractions in real time. This allows practitioners to detect capability-boundary erosion early (the paper shows it emerges around 80 steps) and intervene before irreversible forgetting occurs, rather than relying on final aggregate metrics that mask the trade-off.

  5. Develop a teacher-consistency filter for distillation data. Before applying OPD, filter the teacher-generated reasoning trajectories to include only those that are not already well-supported by the student’s current policy (e.g., using a threshold on the student’s log-probability). The paper shows OPD mainly reinforces existing patterns; by explicitly selecting unfamiliar teacher trajectories, the student is forced to learn genuinely new reasoning structures, potentially converting OPD from illusory to true capability transfer.

What the improved AI system can do:

  • Maintain or expand its maximum reasoning capability (pass@1024) while still improving average performance and low-budget accuracy, avoiding the silent degradation of hard-problem solving.

  • Adaptively allocate inference compute — use fast, distilled models for easy/medium queries and switch to the original model for hard queries, ensuring no loss of ceiling performance.

  • Provide early warnings during training if the model is about to forget rare solutions, allowing automatic rollback or curriculum adjustment.

  • Learn genuinely new reasoning patterns from the teacher rather than merely re-weighting existing ones, leading to actual capability expansion in domains where the student was previously weak.

  • Achieve both efficiency and robustness — faster responses for typical tasks without sacrificing the ability to solve extremely difficult problems that require extensive sampling.

Sources

Related papers