I-SDPO: Instance-Level Adaptive Self-Distillation Policy Optimization

arXiv:2608.12957 · cs.LG, cs.CL · Submitted 2026-08-13 · Read on arXiv

Yubo Zhang, Xinhong Ma, Zezhong Tan, Ziqiang Dong

Qwen Large Model Application Team, Alibaba

cs.LG, cs.CL

Submitted: 2026-08-13

Updated: 2026-08-14

Comments: 14 pages, 3 figures

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 75/100

The gist: I-SDPO: Instance-Level Adaptive Self-Distillation Policy Optimization Summary This paper introduces I-SDPO (Instance-Level Adaptive Self-Distillation Policy Optimization), a method for reinforcement

Terminology

Summary

I-SDPO: Instance-Level Adaptive Self-Distillation Policy Optimization

Summary

This paper introduces I-SDPO (Instance-Level Adaptive Self-Distillation Policy Optimization), a method for reinforcement learning (RL) post-training of large language models (LLMs) that addresses the degenerate gradient problem in Group Relative Policy Optimization (GRPO).

Problem: GRPO computes advantages relative to other samples in a rollout group. When all K sampled responses for a prompt are incorrect, the rewards are similar (r1 ≈ r2 ≈... ≈ rK ≈ 0), causing advantages to collapse to near-zero (Ai = ri − r̄ ≈ 0), producing negligible policy gradients. This is severe during early training or on challenging problems. Privileged self-distillation can fill this gap with dense token supervision, but applying it throughout training creates a different failure mode: the teacher is a biased, low-variance surrogate for the reward objective, so persistent imitation can oppose reward-improving updates after the policy becomes capable of producing successful trajectories.

Method: I-SDPO treats teacher reliance as capability-dependent, making one routing decision per input instance and sharing it across that instance's rollout group. The routing logic is:

  • If any response in a group is correct (ci = 1): all samples go to GRPO, preserving reward contrast in mixed groups.

  • If all responses are wrong and ground-truth exists (ci = 0, mi = 1): all samples go to SDPO (self-distillation).

  • If all responses are wrong but no ground-truth (ci = 0, mi = 0): defaults to GRPO.

The combined loss is: LI-SDPO = Σ(zi GRPO · L GRPO(i) + zi SDPO · L SDPO(i)) / Σ(zi GRPO + zi SDPO)

The SDPO branch uses a privileged teacher that observes the ground-truth solution (prepended to the prompt context), maintained as an EMA of the student, with entropy-aware dynamic weighting (wi,t = exp(−β·H t tea) / Σ exp(−β·H s tea)) and forward-reverse KL interpolation (LSDPO = (1−α) KL(πtea ∥πθ) + α KL(πθ ∥πtea), with α=0.5 by default).

Key theoretical contributions:

  1. Token-space alignment criterion: The descent direction for forward-KL is qs − ps (teacher minus student), while an ideal direction would be us − ps (reward-compatible target minus student). Their alignment is Γs = (qs − ps)⊤(us − ps) = ½(∥qs − ps∥2 + ∥us − ps∥2 − ∥qs − us∥2). Teacher supervision is locally helpful only when Γs > 0.

  2. Bias floor (Proposition 1): Near a reward optimum θ⋆, with LR(θ) ≈ ½∥θ − θ⋆∥2 H and LD(θ) ≈ ½∥θ − (θ⋆ + b)∥2 H, minimizing LR + λLD gives θλ = θ⋆ + (λ/(1+λ))b, and LR(θλ) − LR(θ⋆) ≈ (λ2/(2(1+λ)2))∥b∥2 H. This predicts a bias floor whenever teacher mismatch b and effective distillation weight λ both persist.

  3. Self-annealing property (Proposition 2): Under conditionally independent sampling with per-sample success probability pt, the expected SDPO routing probability is f(t) = (1 − pt) K, which is non-increasing as pt increases. This automatically withdraws teacher influence without a hand-designed schedule.

Experimental results: On SciKnowEval (four scientific domains: biology, material science, chemistry, physics) with Qwen3-8B, after 2 epochs of training with 16 rollout samples per prompt:

Method Biology Material Chemistry Physics Avg.


GRPO 32.12 70.74 62.92 60.88 56.67

SDPO 45.93 71.41 76.64 68.98 65.74

SRPO (sample-level) 44.27 72.94 78.69 68.12 66.01

I-SDPO (instance-level) 50.25 74.53 81.16 75.31 70.31

I-SDPO achieves the best result in all four domains, improving average mean@16 accuracy from 56.67% (GRPO) to 70.31%, with a maximum domain gain of 18.24 points (chemistry). It outperforms GRPO by 13.64 points, pure SDPO by 4.57 points, and sample-level SRPO by 4.30 points on average.

Training dynamics: Distillation-based methods improve rapidly early in training when reward-bearing rollouts are scarce. The GRPO fraction trends upward during training (from 0.65 to 0.90 on biology, 0.78 to 0.96 on physics), while all-wrong group fractions decrease correspondingly, confirming the self-annealing property. I-SDPO assigns a larger fraction to GRPO on moderate-difficulty data (material science and physics, 0.90–0.96 versus SRPO's 0.82–0.90), arguing against more teacher supervision is always better.

Ablations: The balanced forward-reverse KL objective (α=0.5) achieves the highest final accuracy in all four domains (68.16% cross-domain average) compared to forward KL alone (60.39%) and reverse KL alone (61.63%). Distillation top-k sensitivity shows neither routing method improves monotonically with top-k; k=100 is used in main experiments.

Limitations: Experiments use only Qwen3-8B, and evaluation is limited to SciKnowEval. Results may differ on other model families/scales or general mathematical/open-ended reasoning tasks with different reward structures and solution multiplicity.

Improvements for AI systems

Based on the paper, here are the specific improvements I can make to AI systems:

1. Adaptive Teacher-Reliance Routing

  • Implement a capability-dependent routing mechanism that dynamically decides between pure reward optimization (GRPO) and teacher-guided distillation (SDPO) per input instance, rather than using a fixed training schedule.

  • The system automatically withdraws teacher influence as its own success rate increases, eliminating the need for hand-tuned annealing schedules.

2. Bias-Floor-Aware Training

  • Add a theoretical safeguard that detects when persistent teacher mismatch (bias floor) would cause the model to converge to a suboptimal reward optimum, and automatically reduces distillation weight when the policy approaches reward plateaus.

3. Token-Space Alignment Monitoring

  • Compute the alignment score Γs between teacher supervision direction and reward-compatible direction at each training step, and only apply teacher gradients when Γs > 0, preventing harmful imitation that opposes reward improvement.

4. Entropy-Weighted Multi-Teacher Distillation

  • Maintain an ensemble of EMA teachers with different update speeds, and weight their contributions dynamically based on the entropy of their output distributions—giving more weight to confident teachers on easy instances and less on uncertain ones.

5. Balanced Forward-Reverse KL Objective

  • Replace pure forward or reverse KL distillation with an interpolated objective (α=0.5) that balances mode-seeking (reverse KL) and mode-covering (forward KL) behaviors, improving final accuracy by 8 points over either alone.

  • Solve harder reasoning problems earlier in training: By routing all-wrong groups to teacher distillation, the system gets dense token-level supervision when reward signals are absent, accelerating early learning on challenging prompts.

  • Avoid reward hacking and overfitting to teacher bias: The instance-level routing prevents the teacher from dominating once the model becomes competent, preserving reward contrast in mixed-quality groups.

  • Self-adapt to task difficulty: On easy tasks, it quickly transitions to pure reward optimization (GRPO fraction >0.9); on hard tasks, it retains teacher guidance longer, automatically balancing exploration and exploitation.

  • Achieve higher final accuracy: Based on SciKnowEval results, the system improves mean@16 accuracy from 56.7% (GRPO) to 70.3%, with gains up to 18.2 points in chemistry—outperforming both pure GRPO and pure SDPO by 13.6 and 4.6 points respectively.

  • Maintain stable training dynamics: The self-annealing property ensures teacher influence decreases monotonically with model capability, preventing the bias floor failure mode where persistent imitation blocks further reward improvement.

  • Generalize across scientific domains: The method shows consistent improvements across biology, material science, chemistry, and physics, suggesting it can handle diverse reasoning tasks with varying reward densities.

Sources

Related papers