SeOPD: Self-Evolving LLMs via Online Policy Distillation from Self-Generated Chain-of-Thought
cs.AI, cs.CL
Submitted: 2026-09-27
Updated: 2026-10-07
Terminology
Sources
- Putting on the Thinking Hats: A Survey on Chain of Thought Fine-tuning from the Perspective of Human Reasoning Mechanism
- HDPO: Hybrid Distillation Policy Optimization via Privileged Self-Distillation
- SciKnowEval: Evaluating Multi-level Scientific Knowledge of Large Language Models
- Distilling the Knowledge in a Neural Network
- Skill-Conditioned Gated Self-Distillation for LLM Reasoning
- On-Policy Self-Distillation without Any Supervision
- When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale
- OpenAI o1 System Card
- Privileged Information Distillation for Language Models
- Proximal Policy Optimization Algorithms
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- A Survey of On-Policy Distillation for Large Language Models
- ToolAlpaca: Generalized Tool Learning for Language Models with 3000 Simulated Cases
- Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents
- DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
- Qwen3 Technical Report
- Qwen2.5 Technical Report
- On-Policy Context Distillation for Language Models
- On-Policy Distillation with Best-of-N Teacher Rollout Selection
- Instruction-Following Evaluation for Large Language Models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection