Skill-Conditioned Gated Self-Distillation for LLM Reasoning
cs.CL, cs.AI
Submitted: 2026-05-27
Updated: 2026-09-03
Comments: Accepted by EMNLP 2026 Findings. Code is available at https://github.com/walawalagoose/SGSD
Code: https://github.com/walawalagoose/SGSD
License: http://creativecommons.org/licenses/by/4.0/
The gist: On-policy self-distillation (SD) improves LLM reasoning by using teacher-side privileged information (PI) to turn sparse verifier outcomes into dense token-level supervision.
Terminology
Abstract
On-policy self-distillation (SD) improves LLM reasoning by using teacher-side privileged information (PI) to turn sparse verifier outcomes into dense token-level supervision. Existing methods usually assume trusted PI, such as reference answers or successful traces. We ask whether PI can instead come from an experience-derived skill bank, where retrieved skills are compact and reusable but may also be irrelevant or misleading. We propose Skill-Conditioned Gated Self-Distillation (SGSD), which formulates skill-based SD as teacher hypothesis validation rather than unconditional imitation. SGSD retrieves skill-mistake pairs, constructs a multi-teacher pool, and lets all skill-conditioned teachers score the same plain-prompt student rollout. The verifier validates each teacher's polarity: supporting a success or suppressing a failure gives positive supervision, while the opposite stance is reversed. A robust gated objective then distills informative teacher-student disagreements while suppressing uncertain or extreme signals. Experiments on multiple mathematical reasoning benchmarks show that SGSD consistently improves over GRPO and remains competitive with answer-conditioned OPSD under a weaker PI assumption. For example, on Qwen3-1.7B, SGSD outperforms GRPO by 6.2% and OPSD by 1.7% on average on AIME24, AIME25, and HMMT25.
Sources
- Test-Time Distillation for Continual Model Adaptation
- A Brief Overview: On-Policy Self-Distillation In Large Language Models
- Inference-Time Budget Control for LLM Search Agents
- Self-Distillation Zero: Self-Revision Turns Binary Rewards into Dense Supervision
- Distilling the Knowledge in a Neural Network
- Self-Induced Outcome Potential: Turn-Level Credit Assignment for Agents without Verifiers
- Optimizing Agentic Reasoning with Retrieval via Synthetic Semantic Information Gain Reward
- AgentsCoDriver: Large Language Model Empowered Collaborative Driving with Lifelong Learning
- Reinforcement Learning via Self-Distillation
- Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents
- UniSD: Towards a Unified Self-Distillation Framework for Large Language Models
- Reinforcement Learning for Self-Improving Agent with Skill Library
- Unifying Group-Relative and Self-Distillation Policy Optimization via Sample Routing
- Trace2Skill: Distill Trajectory-Local Lessons into Transferable Agent Skills
- SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning
- From Exploration to Exploitation: A Two-Stage Entropy RLVR Approach for Noise-Tolerant MLLM Training
- Qwen3 Technical Report
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Self-Distilled RLVR
- Self-Distillation Enables Continual Learning
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering