Privileged, but Biased: How PI-Conditioned Teachers Break Self-Distillation
Sarthak Harne, Chinmay Karkar, Yash Pandya, Ahmed Awadallah, Akshay Nambi
cs.AI, cs.LG
Submitted: 2026-08-05
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Self-distillation (SD) has emerged as a compute-efficient alternative to reinforcement learning with verifiable rewards: a self-teacher, conditioned on privileged information (PI) about the answer
Terminology
Abstract
Self-distillation (SD) has emerged as a compute-efficient alternative to reinforcement learning with verifiable rewards: a self-teacher, conditioned on privileged information (PI) about the answer such as a reference solution, supplies dense per-token supervision to a student that never sees it. Reported gains, however, come almost exclusively from narrow, low-difficulty settings, leaving open a basic question: as a lone objective, with no reward term, does SD teach anything? We reproduce SDPO's reported gains in its easy setting, then apply the identical setup to difficult tasks and find that it does not. Across question answering, mathematics, coding, and multi-turn agentic tool use, across reasoning modes, model sizes, and forms of PI, and under both the SDPO and OPSD recipes, the per-token loss falls steadily while validation accuracy does not improve and typically degrades. We explain this failure through a single causal chain from the loss to the model it produces. The chain begins with PI bias: having seen one particular reference solution, the teacher's per-token target is pulled toward that trajectory rather than toward correctness in general, an effect we quantify with a PI Bias Score. Trained to match this target everywhere, the student's objective becomes nearly blind to whether a rollout is correct, and the loss it assigns falls mostly on low-information tokens like stopwords, punctuation, uncertainty markers, rather than those that determine the answer; within correct rollouts the exploratory tokens incur the highest divergence, so it penalizes the hesitation that reasoning requires. The result is a flatter, less decisive student that is no better at reasoning: as a lone objective, SD optimizes a signal decoupled from task success.
Sources
- Program Synthesis with Large Language Models
- Evaluating Large Language Models Trained on Code
- SciKnowEval: Evaluating Multi-level Scientific Knowledge of Large Language Models
- Reinforcement Learning via Self-Distillation
- Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?
- Tulu 3: Pushing Frontiers in Open Language Model Post-Training
- Self-Distilled Agentic Reinforcement Learning
- CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Self-Distillation Enables Continual Learning
- Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents
- Self-Distilled RLVR
- DAPO: An Open-Source LLM Reinforcement Learning System at Scale
- Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection