Mismatch Matters: On-Policy Distillation Beyond Token Agreement

arXiv:2608.09836 · cs.AI, cs.CL · Submitted 2026-08-10 · Read on arXiv

Zichao Yu, Chengzhi Yu, Shengze Xu, Yujin Han, Bingqing Jiang, Xu Wang, Difan Zou

The University of Hong Kong · University of Science and Technology of China · The Chinese University of Hong Kong

cs.AI, cs.CL

Submitted: 2026-08-10

Updated: 2026-08-11

Code: https://github.com/yzc-666/TIDE

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 78/100

The gist: Mismatch Matters: On-Policy Distillation Beyond Token Agreement identifies a failure mode in on-policy distillation (OPD) called "degenerate agreement," where students exploit repetitive loops to

Terminology

Summary

Mismatch Matters: On-Policy Distillation Beyond Token Agreement identifies a failure mode in on-policy distillation (OPD) called degenerate agreement, where students exploit repetitive loops to achieve near-perfect token agreement with the teacher despite globally flawed responses. The paper shifts focus from agreement to teacher–student mismatch, categorizing mismatch tokens into two types: student-excess tokens (generated by the student but assigned near-zero probability by the teacher, causing unbounded log-ratio corrections that destabilize updates) and student-deficit tokens (preferred by the teacher but rarely sampled by the student, blocking transfer of the teacher's reasoning patterns). To address these, the authors propose TIDE (Token-level Independent Deficit–Excess correction), which applies bounded Hellinger shaping to suppress severe sampled excesses and an analytic teacher top-K injection to restore deficient probability mass without requiring deficit tokens to be sampled. Across mathematical reasoning benchmarks with multiple Qwen3 teacher–student pairs, TIDE consistently outperforms standard OPD and recent token-selection and reward-shaping baselines. The gains are more pronounced under strong teacher–student mismatch, where TIDE improves Avg@8 from 6.9% to 20.3%, reduces average response length by a factor of 3.6, and substantially reduces formatting failures.

The paper first establishes the problem setup, defining the OPD objective as minimizing expected trajectory-level reverse KL divergence. It then demonstrates student-induced teacher hacking: the student can reduce measured teacher–student divergence by generating contexts that make the teacher's outputs more aligned with its own, without improving reasoning ability. Repetitive generation emerges as an exploit, with the fraction of rollouts containing repetitive loops rising from 16.8% to 48.4% during OPD training. Increasing repetition count simultaneously reduces student–teacher KL divergence and teacher entropy, with sixteen repetitions lowering KL by 63×. The paper shows that supervising only the 60% best-matched positions raises Avg@8 from 6.87 to just 7.18, whereas supervising the 20% most mismatched positions more than doubles it to 14.58.

The directional asymmetry of reverse KL is central: student-excess tokens are readily observable through sampling but can induce unstable updates, while student-deficit tokens are difficult to observe and require analytic teacher-top-K guidance. For student-excess tokens, TIDE replaces the unbounded log-ratio with a Hellinger-shaped weight derived from the squared Hellinger divergence, which enforces bounded suppression without heuristic clipping. The paper proves that this transformation is bounded (between-2 and 0), locally faithful (h(a) ≈ a near agreement), and arises naturally from a proper divergence. For student-deficit tokens, TIDE introduces an analytic teacher top-K objective to restore missing probability mass directly, bypassing the need for rare rollouts. The deficit score quantifies how well the student covers the teacher's preferred continuations, decomposing into a KL term for misallocation within the top-K set and a coverage gap term for student mass falling outside this set.

The joint objective combines excess suppression and deficit recovery, with each branch contributing only at positions selected by its corresponding gate. Experiments use two teacher–student pairs: JustRL-DeepSeek-1.5B → DeepSeek-R1-Distill-Qwen-1.5B (weak mismatch) and Qwen3-8B → Qwen3-1.7B-Base (strong mismatch). Under weak mismatch, TIDE achieves the highest overall Avg@8 and Pass@8 of 46.7% and 65.0%, improving over OPD by 1.0 and 3.2 points. Under strong mismatch, standard OPD falls to the original student's Avg@8 level of 6.9%, whereas TIDE reaches 20.3%, outperforming OPD, FiRe-OPD, GRPO, PowerOPD, and AOPD by 13.4, 5.4, 4.0, 4.6, and 2.0 points respectively. For Pass@8, TIDE obtains 34.2%, compared with 25.3% for OPD, 27.2% for FiRe-OPD, 29.5% for GRPO, 31.2% for PowerOPD, and 34.0% for AOPD.

Generation diagnostics show that under strong mismatch, OPD and FiRe-OPD produce 22,395- and 28,307-token responses, with 65.5% and 54.0% missing a boxed answer; PowerOPD reduces but does not eliminate this degeneration, averaging 13,972 tokens with a 26.3% missing-answer rate. TIDE achieves the highest Avg@8 of 20.3% while reducing response length to 7,294 tokens and the missing-answer rate to 5.4%. This gain cannot be attributed to shorter responses alone: GRPO and AOPD generate even shorter outputs but remain 4.0 and 2.0 Avg@8 points behind TIDE.

Component ablations show that selection alone raises Avg@8 from 6.87 to 14.58 but increases response length from 22.4K to 29.7K tokens. Applying bounded Hellinger shaping further improves Avg@8 to 18.57 and partially controls length increase, reducing it to 23.6K. Analytic deficit recovery independently reaches a similar Avg@8 of 18.32 while producing substantially shorter responses of 13.9K tokens. Combining all three components achieves the best Avg@8 and Pass@8 of 20.34 and 34.17, while further reducing response length to 7.3K tokens. Sensitivity analysis shows that reducing the deficit-recovery weight λ from 1.0 to 0.25 increases response length from 7.3K to 26.9K tokens and format errors from 706 to 5,510, while increasing λ to 2.0 further shortens responses but reduces Avg@8 to 19.20.

The paper concludes that reliable on-policy distillation should allocate supervision according to both the direction and accessibility of disagreement. TIDE assumes a locally reliable teacher, and the evaluation is limited to two teacher–student pairs and mathematical reasoning; dialogue, code, and multilingual settings remain open. The fixed-state, operator-level theory gives no global convergence guarantee, and adaptive keep-rate schedules warrant study.

Improvements for AI systems

Improvements to AI Systems:

  1. Implement TIDE-based distillation loss for student LLMs

Replace standard on-policy distillation (OPD) with the TIDE objective: use bounded Hellinger shaping for student-excess tokens (suppressing unstable, near-zero teacher-probability outputs) and analytic teacher top-K injection for student-deficit tokens (restoring rare but teacher-preferred reasoning paths without sampling). This directly prevents degenerate agreement (repetitive loops) and improves reasoning quality under teacher–student mismatch.

  1. Add a mismatch-aware token gate for supervision allocation

Instead of supervising all tokens or only well-matched ones, dynamically select positions based on teacher–student mismatch direction (excess vs. deficit). This focuses training on the 20% most mismatched tokens, which more than doubles Avg@8 (from 6.87 to 14.58) compared to best-matched-only supervision.

  1. Integrate a bounded divergence correction to stabilize training

Use the Hellinger-shaped weight (proven bounded between-2 and 0, locally faithful near agreement) to replace unbounded log-ratio corrections. This prevents gradient explosions from student-excess tokens, reducing training instability and response degeneration (e.g., cutting response length from 22.4K to 7.3K tokens and format errors from 5,510 to 706).

  1. Add an analytic teacher top-K coverage term for rare reasoning patterns

For student-deficit tokens, compute a deficit score that decomposes into (a) KL misallocation within the teacher’s top-K continuations and (b) a coverage gap for student mass outside that set. Inject this as a direct loss term, eliminating the need to sample rare teacher-preferred tokens—this alone reduces response length by 40% (from 22.4K to 13.9K tokens) while maintaining high accuracy.

  1. Deploy a combined excess–deficit correction with tunable weight

The full TIDE objective (excess suppression + deficit recovery) achieves the best trade-off: Avg@8 of 20.34% and Pass@8 of 34.17% under strong mismatch, with response length reduced by 3.6× and missing-answer rate down to 5.4% (vs. 65.5% for OPD). The weight λ for deficit recovery can be tuned to control response length vs. accuracy (e.g., λ=1.0 optimal; λ=0.25 increases length to 26.9K tokens).

  1. Enable robust distillation for heterogeneous teacher–student pairs

The method works across weak mismatch (e.g., JustRL-DeepSeek-1.5B → DeepSeek-R1-Distill-Qwen-1.5B) and strong mismatch (Qwen3-8B → Qwen3-1.7B-Base), improving Avg@8 by 1.0 and 13.4 points respectively. This allows distilling large, capable teachers into much smaller students without catastrophic performance collapse.

What the improved AI system can do:

  • Generate longer, coherent, and correct reasoning chains without falling into repetitive loops, even when the teacher and student have very different capabilities or token distributions.

  • Learn from a teacher’s rare but valuable reasoning patterns (e.g., multi-step derivations, alternative solution paths) that standard sampling-based distillation would miss, improving pass@k and average performance on math benchmarks.

  • Stabilize training by avoiding unbounded gradient updates from outlier tokens, leading to faster convergence and fewer training failures (e.g., format errors, truncated outputs).

  • Produce concise, high-quality outputs—reducing average response length by up to 3.6× while improving accuracy, making it suitable for latency-sensitive or cost-constrained deployment.

  • Transfer reasoning skills across model families and sizes (e.g., from Qwen3-8B to Qwen3-1.7B-Base) with minimal performance loss, enabling efficient on-device or edge deployment of strong reasoning capabilities.

Sources

Related papers