MAIGO: Mitigating Lost-in-Conversation with History-Cleaned On-Policy Self-Distillation
cs.CL
Submitted: 2026-05-26
Updated: 2026-08-30
Terminology
Sources
- Training Verifiers to Solve Math Word Problems
- When Attention Closes: How LLMs Lose the Thread in Multi-Turn Interaction
- The Llama 3 Herd of Models
- Grounded Continuation: A Linear-Time Runtime Verifier for LLM Conversations
- Models Recall What They Violate: Constraint Adherence in Multi-Turn LLM Ideation
- LLMs Get Lost In Multi-Turn Conversation
- Mitigating Lost in Multi-turn Conversation via Curriculum RL with Verifiable Accuracy and Abstention Rewards
- Evaluating Temporal Consistency in Multi-Turn Language Models
- DeepSeek-V3 Technical Report
- Intent Mismatch Causes LLMs to Get Lost in Multi-Turn Conversation
- SEQUOR: A Multi-Turn Benchmark for Realistic Constraint Following
- Evaluating Large Language Models Trained on Code
- Breaking Contextual Inertia: Reinforcement Learning with Single-Turn Anchors for Stable Multi-Turn Interaction
- Multi-Rollout On-Policy Distillation via Peer Successes and Failures
- Pause or Fabricate? Training Language Models for Grounded Reasoning
- OPSDL: On-Policy Self-Distillation for Long-Context Language Models
- Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- SCOPE: Signal-Calibrated On-Policy Distillation Enhancement with Dual-Path Adaptive Weighting
- Mitigating Conversational Inertia in Multi-Turn Agents
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering