AdviSD: Learning to Advise Frontier LLMs via Targeted Multi-Turn Self-Distillation
cs.AI, cs.CL, cs.LG
Submitted: 2026-09-29
Updated: 2026-09-29
Code: https://github.com/lasgroup/SDPO
Terminology
Sources
- Reinforcement Learning from Rich Feedback with Distributional DAgger
- ACEBench: Who Wins the Match Point in Tool Usage?
- The Trace Is the State: Exact Credit Assignment for LLM Agent Teams
- Distill Where You Fail: Recovering Learning Signals of Negative RL-Groups from Adaptive Teacher Guidance
- Counterfactual Credit Policy Optimization for Multi-Agent Collaboration
- HERO: Hindsight-Enhanced Reflection from Environment Observations for Agentic Self-Distillation
- Unifying distillation and privileged information
- RLCSD: Reinforcement Learning with Contrastive On-Policy Self-Distillation
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment
- TRACE: Distilling Where It Matters via Token-Routed Self On-Policy Alignment
- Flux-OPD: On-Policy Distillation with Evolving Contexts
- Qwen3 Technical Report
- Self-Distilled RLVR
- Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation
- HINT-SD: Targeted Hindsight Self-Distillation for Long-Horizon Agents
- Latent On-Policy Self-Distillation
- SAGE-OPD: Selective Agent-Guided Intervention for Multi-Turn On-Policy Distillation
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection