What Should a Self-Teacher See? Privileged Context Design for On-Policy Self-Distillation
cs.LG, cs.AI
Submitted: 2026-09-22
Updated: 2026-09-26
License: http://creativecommons.org/licenses/by/4.0/
The gist: More privileged information does not always make a better teacher.
Terminology
Abstract
More privileged information does not always make a better teacher. We study this tension in on-policy self-distillation (OPSD), where a frozen copy of the base model scores the student's own rollouts under privileged context, conventionally a complete reference solution that bundles the final answer with one particular reasoning path. Holding the student view and training fixed within each scale, we compare that default against three abstractions compiled offline, a named strategy, a method-independent framing, and a problem category, and against an answer-only control that keeps the destination but removes the path. In the primary runs on competition mathematics, the best intermediate contexts improve the in-domain peak mean over the full solution by 1.4 points at 4B and 1.6 at 8B, while storing an order of magnitude fewer hint tokens. Comparisons across three seeds also show positive mean gains for the framing and category contexts at both scales. Answer-only conditioning remains competitive in the primary runs, within 0.2 points of the full solution at these scales. The preferred context varies with student scale and task. Initial teacher-student KL does not order downstream performance. What a self-teacher should see is therefore not everything it could, but the level of abstraction its student can still act on.
Sources
- Unmasking On-Policy Distillation: Where It Helps, Where It Hurts, and Why
- OpenThoughts: Data Recipes for Reasoning Models
- Adaptive Teacher Exposure for Self-Distillation in LLM Reasoning
- Privileged, but Biased: How PI-Conditioned Teachers Break Self-Distillation
- Distilling the Knowledge in a Neural Network
- Uni-OPD: Unifying On-Policy Distillation with a Dual-Perspective Recipe
- LoRA: Low-Rank Adaptation of Large Language Models
- Privileged Solutions or Context-Induced Teacher Behavior? Dissecting On-Policy Self-Distillation
- Entropy-Aware On-Policy Distillation of Language Models
- Rethinking On-Policy Self-Distillation for Thinking Models
- Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
- On-Policy Self-Distillation without Any Supervision
- ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning
- Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models
- Privileged Information Distillation for Language Models
- Qwen3 Technical Report
- GPQA: A Graduate-Level Google-Proof Q&A Benchmark
- CRISP: Compressed Reasoning via Iterative Self-Policy Distillation
- Self-Distillation Enables Continual Learning
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks