Privileged Solutions or Context-Induced Teacher Behavior? Dissecting On-Policy Self-Distillation
Yuki Ichihara, Naoto Iwase, Mohammad Atif Quamar, Junpei Komiyama
Mohamed bin Zayed University of Artificial Intelligence · Nagoya University · RIKEN AIP
cs.LG, cs.AI
Submitted: 2026-08-10
Updated: 2026-08-11
Comments: Working progress
Code: https://github.com/MBZUAI-reasoninglab/OP2SD
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 75/100
The gist: On-Policy Self-Distillation (OPSD) is commonly interpreted as the transfer of privileged information: a teacher observes the verified solution to the target problem and supervises the student’s
Terminology
Summary
On-Policy Self-Distillation (OPSD) is commonly interpreted as the transfer of privileged information: a teacher observes the verified solution to the target problem and supervises the student’s trajectory. However, this interpretation conflates two effects. The reference solution not only reveals the answer to the current instance but also changes the context under which the teacher provides token-level supervision. We investigate the role of target-specific privilege with OP2SD (On-Policy Self-Distillation from Other Problems), which replaces the paired reference with a problem and solution from a different example, while preserving the student rollout, teacher, and distillation objective. Across three models and three mathematics benchmarks, OP2SD improves over the base model, remains competitive with OPSD. The success of OP2SD implies that OPSD gains do not necessarily come from access to the reference solution, and that the teacher’s context-induced behavior is an important factor.
The paper introduces OP2SD to probe whether OPSD depends on the identity of the solution or whether part of its effect arises from behavior induced by the teacher’s additional context. For a target problem A, standard OPSD gives the teacher the verified solution to A. OP2SD instead gives it a different problem B and its solution, while explicitly stating that the example is neither a solution nor a hint for A. The student still sees only A. The intervention breaks the target and reference pairing while retaining the teacher-only worked-solution context. Across the evaluated Qwen3-1.7B, Qwen3-4B, and Qwen3-8B non-thinking models, OP2SD improves accuracy over the corresponding base model on AIME 2024, AIME 2025, and HMMT 2025. Therefore, the paired derivation and answer are not necessary for obtaining the OPSD-like improvements observed in these experiments.
The main results show that for Qwen3-1.7B, OPSD and OP2SD improve Avg@12 over Base on all three benchmarks. For Qwen3-4B and Qwen3-8B, OP2SD achieves the highest Avg@12 point estimate on all three benchmarks for both model sizes. For Qwen3-8B, it outperforms OPSD in Avg@12 by 10.83, 10.69, and 9.38 points on AIME 2024, AIME 2025, and HMMT 2025, respectively. The paper also examines accuracy under fixed token budgets, showing different scaling behavior: for Qwen3-4B, OP2SD remains above OPSD over nearly the entire budget range and reaches a higher final accuracy with fewer mean generated tokens; for Qwen3-8B, the two methods are comparable at small budgets, after which OP2SD continues to improve while OPSD saturates.
The paper further investigates which properties of the teacher signal account for the results. It disentangles the effect of solution conditioning from that of the student-teacher thinking-mode asymmetry. For Qwen3-1.7B, the student generates in non-thinking mode while the teacher scores in thinking mode. Two controls are introduced: Target-only (teacher receives only the problem) and Answer-only (teacher additionally receives the final answer but not its derivation). In the 1.7B setting, both controls improve the base model’s accuracy across all three benchmarks, indicating that improvements can arise even when the teacher receives neither a worked solution nor additional mathematical context. However, in the Qwen3-4B non-thinking setting, Target-only performs substantially below Base on all three benchmarks, and Answer-only partially recovers performance but remains below Base. In this mode-matched configuration, improvement is observed only when the teacher is provided with a complete mathematical worked solution.
The paper then examines what properties of the other-problem context matter. It tests whether diversity among worked examples is necessary by conditioning the teacher on a single fixed algebra problem throughout training. The fixed-example run attains Avg@12 point estimates of 35.83, 31.94, and 17.71 on AIME 2024, AIME 2025, and HMMT 2025, compared with 31.53, 30.62, and 16.11 for varying-example OP2SD with Qwen3-4B. Thus, an OPSD-like improvement can occur even when the auxiliary context is reduced to one repeated example. The paper also holds the fixed worked example constant while varying its solution: a locally corrupted solution performs comparably to or better than the concise correct solution, whereas a substantially more verbose correct solution yields lower Avg@12 and shorter student responses. A fixed trivial 1+1 example performs poorly, achieving Avg@12 scores of only 5.21, 5.28, and 2.29 on AIME 2024, AIME 2025, and HMMT 2025, respectively.
The paper also tests whether broad domain matching is necessary. Using Omni-MATH, it trains on 1,280 Algebra problems and evaluates on 30 held-out Algebra problems. The teacher receives either another Algebra problem and its solution or a Geometry problem and its solution. The Geometry condition is 1.67 points higher, but this apparent advantage is driven by one target near the boundary between Algebra and analytic Geometry. Excluding that problem reverses the ordering: Algebra obtains 59.41% and Geometry obtains 58.98% Avg@12. The paper finds no clear evidence that matching a coarse domain label is necessary.
Finally, the paper tests whether mathematical context is necessary by replacing the teacher’s mathematical worked examples with CAMEL Physics problem and solution pairs. Replacing mathematical worked solutions with physics examples reduces Avg@12 from 31.53 to 18.68 on AIME 2024, from 30.62 to 20.21 on AIME 2025, and from 16.11 to 8.68 on HMMT 2025. The physics condition is also below Base on all three benchmarks. The accuracy drop is accompanied by substantially less reliable answer formatting. Together, these results suggest that the mathematical solution context, or some property associated with it, is important for obtaining the OP2SD gain.
The paper concludes that across three model settings and three mathematics benchmarks, OP2SD remains competitive with OPSD, showing that paired target solutions are not necessary for OPSD-like gains. Most of the OPSD gain is retained, or even reinforced, after the reference solution to the target problem is replaced by the solution to a different problem. The gain persists when the diversity of the problem-solution pair given to the teacher is reduced to a single fixed example. At the same time, trivial mathematical and cross-subject physics contexts fail to preserve the improvement, indicating that arbitrary additional text is insufficient. The verbose contexts also reduced the effectiveness of the teacher. Taken together, these results imply that the major source of the observed gain is not the privileged answer to the target problem, but the change in the teacher’s token-level behavior induced by a mathematical context that works.
Improvements for AI systems
Improvements to AI Systems:
- Decouple privileged information from teacher context in distillation pipelines.
-
What to do: Modify on-policy self-distillation (OPSD) training so that the teacher’s supervision signal is generated using a different problem’s solution (OP2SD) rather than the target’s verified answer.
-
Improved capability: The student model gains accuracy on math benchmarks (e.g., +10.83 points on AIME 2024 for Qwen3-8B) without needing access to the target solution, making the method applicable in privacy-sensitive or answer-unavailable settings.
- Inject a fixed, non-trivial mathematical worked example into the teacher’s context.
-
What to do: Replace the per-instance reference solution with a single, fixed algebra problem-solution pair (e.g., a concise correct algebra derivation) repeated across all training examples.
-
Improved capability: Achieves OPSD-like gains (e.g., 35.83 Avg@12 on AIME 2024 vs. 31.53 for varying examples) while reducing memory and data preparation overhead, enabling more efficient training.
- Use a mode-asymmetric teacher-student configuration for smaller models.
-
What to do: For models ≤1.7B, have the student generate in non-thinking mode while the teacher scores in thinking mode, and provide the teacher only the problem (or problem + answer) without a worked solution.
-
Improved capability: Boosts base accuracy across all three benchmarks (e.g., AIME 2024 improves from 20 to 30 Avg@12) even without any mathematical context, reducing the need for curated solution datasets.
- Filter teacher contexts to avoid verbose or trivial content.
-
What to do: During distillation, exclude teacher contexts that are overly verbose (e.g., multi-page solutions) or trivial (e.g., “1+1=2”), as these degrade student accuracy and response length.
-
Improved capability: Prevents performance collapse (e.g., trivial context drops Avg@12 to 5 on AIME 2024) and maintains consistent answer formatting, improving reliability in production.
- Retain mathematical solution context, not cross-domain text, for teacher conditioning.
-
What to do: Ensure the teacher’s auxiliary context is always a mathematical worked example (e.g., algebra or geometry), not physics or other subjects.
-
Improved capability: Preserves the distillation gain (e.g., 31.53 vs. 18.68 Avg@12 on AIME 2024 when using physics), preventing large accuracy drops and formatting errors.
- Adapt the teacher’s context dynamically based on model size.
-
What to do: For larger models (≥4B), provide the teacher with a full worked solution; for smaller models, use only the problem or answer.
-
Improved capability: Maximizes accuracy across model scales—e.g., Qwen3-4B requires a complete solution to improve, while Qwen3-1.7B improves even with minimal context—allowing a single training framework to handle heterogeneous model sizes.
Sources
- Reinforced Self-Training (ReST) for Language Modeling
- Reinforcement Learning via Self-Distillation
- When Are Teacher Tokens Reliable? Position-Weighted On-Policy Self-Distillation for Reasoning
- Unifying distillation and privileged information
- Semantic Soft Bootstrapping: Long Context Reasoning in LLMs without Reinforcement Learning
- Privileged Information Distillation for Language Models
- Purified OPSD: On-Policy Self-Distillation Without Losing How to Think
- Self-Distillation Enables Continual Learning
- Learning by Distilling Context
- Self-Supervised On-Policy Distillation for Reasoning Language Models
- Gemma 2: Improving Open Language Models at a Practical Size
- Qwen3 Technical Report
- On-Policy Context Distillation for Language Models
- Beyond Imitation: Reflective On-Policy Self-Distillation for LLM Reasoning
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks