Instruct, Not Answer: Using Instruction Privileges in On-Policy Context Distillation
cs.AI, cs.LG
Submitted: 2026-09-26
Updated: 2026-09-26
Terminology
Sources
- A General Language Assistant as a Laboratory for Alignment
- Olmo 3
- Towards Understanding On-Policy Distillation through the Lens of Test-Time Scaling
- Reinforcement Learning via Self-Distillation
- Geometric Self-Distillation for Reasoning Generalization
- Hindsight Compacts but Does Not Repair: Rethinking On-Policy Self-Distillation in Reasoning Models
- Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?
- DemoPSD: Disagreement-Modulated Policy Self-Distillation
- RLCSD: Reinforcement Learning with Contrastive On-Policy Self-Distillation
- DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
- Learning by Distilling Context
- TRACE: Distilling Where It Matters via Token-Routed Self On-Policy Alignment
- Qwen3 Technical Report
- $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
- On-Policy Context Distillation for Language Models
- Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models
- The Many Faces of On-Policy Distillation: Pitfalls, Mechanisms, and Fixes
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection