Mismatch Matters: On-Policy Distillation Beyond Token Agreement
Zichao Yu, Chengzhi Yu, Shengze Xu, Yujin Han, Bingqing Jiang, Xu Wang, Difan Zou
The University of Hong Kong · University of Science and Technology of China · The Chinese University of Hong Kong
cs.AI, cs.CL
Submitted: 2026-08-10
Updated: 2026-08-11
Code: https://github.com/yzc-666/TIDE
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 78/100
The gist: Mismatch Matters: On-Policy Distillation Beyond Token Agreement identifies a failure mode in on-policy distillation (OPD) called "degenerate agreement," where students exploit repetitive loops to
Terminology
Summary
Mismatch Matters: On-Policy Distillation Beyond Token Agreement identifies a failure mode in on-policy distillation (OPD) called degenerate agreement,
where students exploit repetitive loops to achieve near-perfect token agreement with the teacher despite globally flawed responses. The paper shifts focus from agreement to teacher–student mismatch, categorizing mismatch tokens into two types: student-excess tokens (generated by the student but assigned near-zero probability by the teacher, causing unbounded log-ratio corrections that destabilize updates) and student-deficit tokens (preferred by the teacher but rarely sampled by the student, blocking transfer of the teacher's reasoning patterns). To address these, the authors propose TIDE (Token-level Independent Deficit–Excess correction), which applies bounded Hellinger shaping to suppress severe sampled excesses and an analytic teacher top-K injection to restore deficient probability mass without requiring deficit tokens to be sampled. Across mathematical reasoning benchmarks with multiple Qwen3 teacher–student pairs, TIDE consistently outperforms standard OPD and recent token-selection and reward-shaping baselines. The gains are more pronounced under strong teacher–student mismatch, where TIDE improves Avg@8 from 6.9% to 20.3%, reduces average response length by a factor of 3.6, and substantially reduces formatting failures.
The paper first establishes the problem setup, defining the OPD objective as minimizing expected trajectory-level reverse KL divergence. It then demonstrates student-induced teacher hacking
: the student can reduce measured teacher–student divergence by generating contexts that make the teacher's outputs more aligned with its own, without improving reasoning ability. Repetitive generation emerges as an exploit, with the fraction of rollouts containing repetitive loops rising from 16.8% to 48.4% during OPD training. Increasing repetition count simultaneously reduces student–teacher KL divergence and teacher entropy, with sixteen repetitions lowering KL by 63×. The paper shows that supervising only the 60% best-matched positions raises Avg@8 from 6.87 to just 7.18, whereas supervising the 20% most mismatched positions more than doubles it to 14.58.
The directional asymmetry of reverse KL is central: student-excess tokens are readily observable through sampling but can induce unstable updates, while student-deficit tokens are difficult to observe and require analytic teacher-top-K guidance. For student-excess tokens, TIDE replaces the unbounded log-ratio with a Hellinger-shaped weight derived from the squared Hellinger divergence, which enforces bounded suppression without heuristic clipping. The paper proves that this transformation is bounded (between-2 and 0), locally faithful (h(a) ≈ a near agreement), and arises naturally from a proper divergence. For student-deficit tokens, TIDE introduces an analytic teacher top-K objective to restore missing probability mass directly, bypassing the need for rare rollouts. The deficit score quantifies how well the student covers the teacher's preferred continuations, decomposing into a KL term for misallocation within the top-K set and a coverage gap term for student mass falling outside this set.
The joint objective combines excess suppression and deficit recovery, with each branch contributing only at positions selected by its corresponding gate. Experiments use two teacher–student pairs: JustRL-DeepSeek-1.5B → DeepSeek-R1-Distill-Qwen-1.5B (weak mismatch) and Qwen3-8B → Qwen3-1.7B-Base (strong mismatch). Under weak mismatch, TIDE achieves the highest overall Avg@8 and Pass@8 of 46.7% and 65.0%, improving over OPD by 1.0 and 3.2 points. Under strong mismatch, standard OPD falls to the original student's Avg@8 level of 6.9%, whereas TIDE reaches 20.3%, outperforming OPD, FiRe-OPD, GRPO, PowerOPD, and AOPD by 13.4, 5.4, 4.0, 4.6, and 2.0 points respectively. For Pass@8, TIDE obtains 34.2%, compared with 25.3% for OPD, 27.2% for FiRe-OPD, 29.5% for GRPO, 31.2% for PowerOPD, and 34.0% for AOPD.
Generation diagnostics show that under strong mismatch, OPD and FiRe-OPD produce 22,395- and 28,307-token responses, with 65.5% and 54.0% missing a boxed answer; PowerOPD reduces but does not eliminate this degeneration, averaging 13,972 tokens with a 26.3% missing-answer rate. TIDE achieves the highest Avg@8 of 20.3% while reducing response length to 7,294 tokens and the missing-answer rate to 5.4%. This gain cannot be attributed to shorter responses alone: GRPO and AOPD generate even shorter outputs but remain 4.0 and 2.0 Avg@8 points behind TIDE.
Component ablations show that selection alone raises Avg@8 from 6.87 to 14.58 but increases response length from 22.4K to 29.7K tokens. Applying bounded Hellinger shaping further improves Avg@8 to 18.57 and partially controls length increase, reducing it to 23.6K. Analytic deficit recovery independently reaches a similar Avg@8 of 18.32 while producing substantially shorter responses of 13.9K tokens. Combining all three components achieves the best Avg@8 and Pass@8 of 20.34 and 34.17, while further reducing response length to 7.3K tokens. Sensitivity analysis shows that reducing the deficit-recovery weight λ from 1.0 to 0.25 increases response length from 7.3K to 26.9K tokens and format errors from 706 to 5,510, while increasing λ to 2.0 further shortens responses but reduces Avg@8 to 19.20.
The paper concludes that reliable on-policy distillation should allocate supervision according to both the direction and accessibility of disagreement. TIDE assumes a locally reliable teacher, and the evaluation is limited to two teacher–student pairs and mathematical reasoning; dialogue, code, and multilingual settings remain open. The fixed-state, operator-level theory gives no global convergence guarantee, and adaptive keep-rate schedules warrant study.
Improvements for AI systems
Improvements to AI Systems:
- Implement TIDE-based distillation loss for student LLMs
Replace standard on-policy distillation (OPD) with the TIDE objective: use bounded Hellinger shaping for student-excess tokens (suppressing unstable, near-zero teacher-probability outputs) and analytic teacher top-K injection for student-deficit tokens (restoring rare but teacher-preferred reasoning paths without sampling). This directly prevents degenerate agreement (repetitive loops) and improves reasoning quality under teacher–student mismatch.
- Add a mismatch-aware token gate for supervision allocation
Instead of supervising all tokens or only well-matched ones, dynamically select positions based on teacher–student mismatch direction (excess vs. deficit). This focuses training on the 20% most mismatched tokens, which more than doubles Avg@8 (from 6.87 to 14.58) compared to best-matched-only supervision.
- Integrate a bounded divergence correction to stabilize training
Use the Hellinger-shaped weight (proven bounded between-2 and 0, locally faithful near agreement) to replace unbounded log-ratio corrections. This prevents gradient explosions from student-excess tokens, reducing training instability and response degeneration (e.g., cutting response length from 22.4K to 7.3K tokens and format errors from 5,510 to 706).
- Add an analytic teacher top-K coverage term for rare reasoning patterns
For student-deficit tokens, compute a deficit score that decomposes into (a) KL misallocation within the teacher’s top-K continuations and (b) a coverage gap for student mass outside that set. Inject this as a direct loss term, eliminating the need to sample rare teacher-preferred tokens—this alone reduces response length by 40% (from 22.4K to 13.9K tokens) while maintaining high accuracy.
- Deploy a combined excess–deficit correction with tunable weight
The full TIDE objective (excess suppression + deficit recovery) achieves the best trade-off: Avg@8 of 20.34% and Pass@8 of 34.17% under strong mismatch, with response length reduced by 3.6× and missing-answer rate down to 5.4% (vs. 65.5% for OPD). The weight λ for deficit recovery can be tuned to control response length vs. accuracy (e.g., λ=1.0 optimal; λ=0.25 increases length to 26.9K tokens).
- Enable robust distillation for heterogeneous teacher–student pairs
The method works across weak mismatch (e.g., JustRL-DeepSeek-1.5B → DeepSeek-R1-Distill-Qwen-1.5B) and strong mismatch (Qwen3-8B → Qwen3-1.7B-Base), improving Avg@8 by 1.0 and 13.4 points respectively. This allows distilling large, capable teachers into much smaller students without catastrophic performance collapse.
What the improved AI system can do:
-
Generate longer, coherent, and correct reasoning chains without falling into repetitive loops, even when the teacher and student have very different capabilities or token distributions.
-
Learn from a teacher’s rare but valuable reasoning patterns (e.g., multi-step derivations, alternative solution paths) that standard sampling-based distillation would miss, improving pass@k and average performance on math benchmarks.
-
Stabilize training by avoiding unbounded gradient updates from outlier tokens, leading to faster convergence and fewer training failures (e.g., format errors, truncated outputs).
-
Produce concise, high-quality outputs—reducing average response length by up to 3.6× while improving accuracy, making it suitable for latency-sensitive or cost-constrained deployment.
-
Transfer reasoning skills across model families and sizes (e.g., from Qwen3-8B to Qwen3-1.7B-Base) with minimal performance loss, enabling efficient on-device or edge deployment of strong reasoning capabilities.
Sources
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- JustRL: Scaling a 1.5B LLM with a Simple RL Recipe
- Asymmetric On-Policy Distillation: Bridging Exploitation and Imitation at the Token Level
- Trajectory-Refined Distillation
- Respecting Self-Uncertainty in On-Policy Self-Distillation for Efficient LLM Reasoning
- Scaling Reasoning Efficiently via Relaxed On-Policy Distillation
- Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
- Filter, Then Reweight: Rethinking Optimization Granularity in On-Policy Distillation
- Multi-Turn On-Policy Distillation with Prefix Replay
- Teacher-Guided Policy Optimization for On-Policy Reasoning Distillation under Large Policy Divergence
- Your Teacher Can't Help You Here: Combating Supervision Fidelity Decay in On-Policy Distillation
- Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models
- KL for a KL: On-Policy Distillation with Control Variate Baseline
- Proximal Policy Optimization Algorithms
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- HybridFlow: A Flexible and Efficient RLHF Framework
- Kimi K3: Open Frontier Intelligence
- Not All Disagreement Is Learnable: Token Teachability in On-Policy Distillation
- MiMo-V2-Flash Technical Report
- Escaping the KL Agreement Trap in On-Policy Distillation
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection