DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models
ZhiYan Hou, Xinyu Tang, Hongyan An, Jianjin Zhang, Weizhen Wang, Yunyun Han, Gengsheng Li, Xiangzhao Hao, Haiyun Guo, Wenbin Hu, Jinqiao Wang, Yafeng Deng
Institute of Automation, Chinese Academy of Sciences · EverMind · Shanda Group · University of Chinese Academy of Sciences · Wuhan AI Research · Wuhan University
cs.AI
Submitted: 2026-08-07
Updated: 2026-08-10
Comments: 17 pages, 4 figures, 9 tables. Code at https://github.com/DBtxy/DASH-OPSD
Code: https://github.com/DBtxy/DASH-OPSD
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 71/100
The gist: The paper addresses a limitation in Reinforcement Learning with Verifiable Rewards (RLVR) and on-policy self-distillation (OPSD).
Terminology
Summary
The paper addresses a limitation in Reinforcement Learning with Verifiable Rewards (RLVR) and on-policy self-distillation (OPSD). While RLVR improves reasoning capabilities using automatically verifiable outcome signals, these signals are typically sparse and at the sequence-level.
OPSD mitigates this sparsity by querying a privileged teacher at student-visited prefixes to provide dense token-level distributional supervision.
However, the authors find that standard OPSD still underexploits the temporal structure of the rollout. It assigns every local divergence the same coefficient, regardless of its position or the divergence sequence in which it occurs.
The key insight is that the same local discrepancy value can arise after different discrepancy histories that reflect how the mismatch between teacher and student has evolved over student-visited prefixes.
Since the local scalar alone cannot distinguish these temporal contexts, vanilla OPSD cannot adapt token-level distillation weights to the temporal evolution of the discrepancy sequence.
The authors propose Divergence-Adaptive Supervision Horizons (DASH), which maps the gap between each local distillation signal and the sequence-level mean to an adaptive propagation gate and then uses these gates to control backward multi-step aggregation.
-
Local distillation signals: DASH computes per-position forward KL divergence from the privileged teacher to the student: d t = D KL(pi t T pi t S), with vocabulary-level contributions clipped at tau = 0.05.
-
Adaptive propagation gates: At each position, the gap between the local signal and the sequence mean g t = r t - is converted into a gate: lambda t = sg[sigma(-kappa g t)], where kappa = 5 controls sensitivity.
-
Backward multi-step aggregation: The gates control a backward recursion A T = r T, A t = r t + lambda t A t+1, producing the objective L DASH = 1 over T sum t=1 T A t.
This yields weighted form L DASH = 1 over T sum k=1 T c k r k, where c k = 1 + lambda k-1c k-1, making the effective supervision horizon adapt to the realized discrepancy sequence.
The paper provides a fixed-horizon gradient decomposition (Proposition 1/2) showing that the exact gradient of the expected on-policy objective contains, beyond the direct distillation term retained by vanilla OPSD, a trajectory score-function term with future-divergence coefficients.
The authors state this decomposition provides only structural motivation: DASH does not estimate G D u+1, introduce score-function gradients, or perform future-to-past credit assignment.
Across three model scales (Qwen3-1.7B, Qwen3-4B, Qwen3-8B) and three mathematical reasoning benchmarks (AIME 2024, AIME 2025, HMMT February 2025), DASH obtains the highest score among the compared results in all nine benchmark–model settings and the highest macro-average at each model scale.
Specifically:
-
Qwen3-1.7B: DASH raises four-seed OPSD average from 41.87 to 45.07 (+3.20)
-
Qwen3-4B: from 63.60 to 65.00 (+1.40)
-
Qwen3-8B: from 64.80 to 66.40 (+1.60)
DASH also improves over the matched OPSD rerun on every benchmark at every model scale.
-
Adaptive vs. fixed propagation: Fixed coefficients lambda in 0.1, 0.3, 0.5, 0.7, 0.9 all improve over OPSD (best fixed: 43.63 at lambda=0.1), but DASH outperforms the best fixed setting by 1.44 points, showing
fixed multi-step aggregation accounts for part of the gain, while adapting the propagation gates to the realized discrepancy sequence provides an additional improvement.
-
Gate direction: Inverse-gap (reversing the sign of the gate function) reaches only 42.10,
supporting the sign used by DASH among the two tested adaptive mappings.
-
Coefficient allocation vs. average scale: A 2×2 factorial comparison shows "dynamic allocation improves performance by 2.40 points at the OPSD scale and 2.50 points at the DASH scale. In contrast, increasing the average scale contributes only 0.70 points under uniform allocation and 0.80 points under dynamic allocation,
demonstrating that
the improvement primarily comes from DASH's discrepancy-conditioned coefficient allocation."
-
**Propagation sensitivity kappa **: All tested values kappa in 1, 2, 5, 10, 20 outperform OPSD; kappa=5 is best (45.07), with
lower performance at both smaller and larger values suggesting that moderate gate sensitivity better balances insufficient adaptation and overly sharp responses.
-
Divergence choice: Forward KL (45.07) outperforms symmetric JSD (38.23) and reverse KL (41.47).
-
Vocabulary support: Top-100 plus tail achieves 44.37 (only 0.70 below full vocabulary), while top-1 support drops to 34.83.
DASH reuses the teacher and student distributions that OPSD already computes, so the gains require no additional teacher or student forward pass.
Its extra scalar backward scan accounts for less than 1% of the step time.
The three stated contributions are: (1) identifying the temporal coefficient allocation gap in vanilla OPSD; (2) proposing DASH, which uses sequence-relative divergence gaps to adapt effective supervision horizons; and (3) demonstrating that DASH obtains the highest overall scores among compared results across three benchmarks and three model scales, supported by comprehensive ablation studies.
The appendix explores compatibility with outcome-level RL by adding a GRPO correctness term to DASH. Aggregation improves all four matched 1.7B settings,
with the strongest hybrid at eta=0.3, K=8 achieving 44.83 on 1.7B (slightly below DASH's 45.07), but at 8B the same hybrid reaches 67.30 versus DASH's 66.40, indicating compatibility is scale-dependent.
Improvements for AI systems
Improvements to AI systems:
- Replace uniform token-level distillation with divergence-adaptive temporal weighting.
Instead of assigning the same weight to every local teacher-student divergence in on-policy self-distillation, compute per-position forward-KL signals d t, compare each to the sequence-level mean, and convert the gap into a gate lambda t = sg[sigma(-kappa(d t -))]. Use these gates in a backward aggregation A T = d T, A t = d t + lambda t A t+1, and optimize L = 1 over T sum t A t. This makes the effective supervision horizon adapt to the realized discrepancy trajectory.
- Add a fixed multi-step propagation baseline before introducing adaptivity.
Even a non-adaptive multi-step aggregation with a constant lambda in 0.1, 0.3, 0.5, 0.7, 0.9 improves over vanilla OPSD, with lambda=0.1 giving a strong baseline. This isolates the benefit of extending supervision horizons from the benefit of conditioning them on the divergence sequence. The final DASH system should include both: fixed propagation for robustness and adaptive gating for extra gain.
- Use gap-signed gates with moderate sensitivity.
Set lambda t = sg[sigma(-5(d t -))]. This causes positions with below-average local divergence to propagate future losses more strongly, while above-average divergence positions stop propagation sooner. Avoid reversing the gate sign, and avoid using overly sharp or overly flat sensitivity; kappa=5 balances adaptation and stability.
- Use forward KL and a sufficiently rich vocabulary support.
Use forward KL rather than symmetric JSD or reverse KL for the local distillation signal. When reducing vocabulary support, keep the top-100 tokens plus the tail rather than only the top-1 token; top-1 support severely degrades performance. This preserves most of the adaptive-distillation benefit while lowering computational cost.
- Combine DASH with outcome-level RL when beneficial.
Add a GRPO correctness term to DASH with a mixing coefficient eta. Use smaller eta and moderate rollout count (e.g., eta=0.3, K=8) for smaller models, and allow the hybrid to scale to larger models where it can exceed pure DASH. This gives a tunable mechanism for trading dense teacher supervision against sparse verifiable rewards.
- Keep the extra cost negligible.
Implement DASH by reusing the teacher and student distributions already computed by OPSD. The backward scalar scan adds less than 1% of step time, so the improved system gains accuracy without requiring additional teacher or student forward passes.
What the improved AI system can do:
-
Outperform vanilla on-policy self-distillation on mathematical reasoning benchmarks across multiple model scales: e.g., +3.20 points on Qwen3-1.7B, +1.40 points on Qwen3-4B, and +1.60 points on Qwen3-8B macro-average.
-
Achieve the highest score among compared methods in all nine benchmark–model settings tested, including AIME 2024, AIME 2025, and HMMT February 2025.
-
Improve over the matched OPSD rerun on every benchmark at every model scale.
-
Dynamically allocate supervision weight across reasoning steps based on how the teacher-student discrepancy evolves, rather than treating each token identically.
-
Provide a principled middle ground between sparse outcome-level RL and dense token-level distillation, enabling better credit assignment for long reasoning chains.
-
Operate at nearly the same training cost as vanilla OPSD while delivering consistent accuracy gains.
Abstract
Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models using automatically verifiable outcome signals, but these signals are typically sparse and at the sequence-level. On-policy self-distillation (OPSD) mitigates this sparsity by querying a privileged teacher at student-visited prefixes and providing dense token-level distributional supervision. Although this dense supervision alleviates signal sparsity, we find that standard OPSD still underexploits the temporal structure of the rollout. It assigns every local divergence the same coefficient, regardless of its position or the divergence sequence in which it occurs. In on-policy autoregressive generation, the same divergence magnitude can follow different discrepancy histories, reflecting different evolutions of the mismatch between the teacher and student. Since the local scalar alone cannot distinguish these temporal contexts, standard OPSD cannot adapt its token-level weights to the realized discrepancy sequence. To address this limitation, we propose Divergence-Adaptive Supervision Horizons (DASH). DASH maps the gap between each local distillation signal and the sequence-level mean to an adaptive propagation gate and then uses these gates to control backward multi-step aggregation. By doing so, DASH adjusts token-level supervision weights according to how local divergences evolve during generation. Experiments on three mathematical reasoning benchmarks across three model scales show that DASH improves over our matched vanilla OPSD reruns on every benchmark at all three scales. DASH reuses the teacher and student distributions that OPSD already computes, so the gains require no additional teacher or student forward pass. Code: https://github.com/DBtxy/DASH-OPSD
Sources
- Evaluating Large Language Models Trained on Code
- Training Verifiers to Solve Math Word Problems
- OpenThoughts: Data Recipes for Reasoning Models
- Distilling the Knowledge in a Neural Network
- Entropy-Aware On-Policy Distillation of Language Models
- VinePPO: Refining Credit Assignment in RL Training of LLMs
- Tulu 3: Pushing Frontiers in Open Language Model Post-Training
- PHF: Privileged Hidden Flow for On-Policy Self-Distillation
- ReNIO: Reweighting Negative Trajectory Importance for LLM On-Policy Distillation
- When Are Teacher Tokens Reliable? Position-Weighted On-Policy Self-Distillation for Reasoning
- Your Teacher Can't Help You Here: Combating Supervision Fidelity Decay in On-Policy Distillation
- AVSD: Adaptive-View Self-Distillation by Balancing Consensus and Teacher-Specific Privileged Signals
- GRPO-$\lambda$: Credit Assignment improves LLM Reasoning
- Proximal Policy Optimization Algorithms
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Purified OPSD: On-Policy Self-Distillation Without Losing How to Think
- Solving math word problems with process- and outcome-based feedback
- On the Position Bias of On-Policy Distillation
- Qwen3 Technical Report
- DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection