REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation
Shanghai Artificial Intelligence Laboratory · Peking University · Southwest Jiaotong University · University of Science and Technology of China · University of Electronic Science and Technology of China · Renmin University of China · Frontier Discovery Center, Shanghai Artificial Intelligence Laboratory · Autonomous Driving, Shanghai Artificial Intelligence Laboratory
cs.LG, cs.AI
Submitted: 2026-08-12
Updated: 2026-08-27
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 75/100
The gist: On-policy distillation (OPD) trains a student on its own generated trajectories under dense token-level supervision from a teacher, providing an effective post-training paradigm for large language
Terminology
Summary
On-policy distillation (OPD) trains a student on its own generated trajectories under dense token-level supervision from a teacher, providing an effective post-training paradigm for large language models. Reward-extrapolation methods such as ExOPD further amplify the teacher–reference log-likelihood ratio to move beyond direct imitation. However, ExOPD uses a single global scalar λ to apply the same extrapolation strength indiscriminately to every token. This can drive the student to aggressively fit extreme peaks in the teacher–reference log-ratio that defines the implicit reward, resulting in reward hacking and unstable training. Moreover, the optimal λ varies across domains, requiring costly domain-specific sweeps that may still fail to identify an appropriate extrapolation strength. We propose REOPD, a Reliability-Adaptive Reward Extrapolation framework for On-Policy Distillation. REOPD combines a token-level compatibility weight with a batch-level adaptive budget. The former modulates token-wise residuals according to the student–teacher discrepancy, while the latter dynamically adjusts the overall extrapolation strength according to the reliability and scale of residual signals in each batch.
REOPD preserves the standard OPD alignment term and adapts only the additional teacher–reference residual. It combines a token-level compatibility weight with a bounded micro-batch budget. Their product defines a token-wise effective coefficient, allowing reliable residuals to be emphasized without uniformly increasing every update. The method reuses the student, teacher, and reference log-probabilities already available in G-OPD and requires no verifier, reward model, value model, or additional rollout.
The token-level compatibility weight is defined as follows. For a sampled token yt, the student–teacher alignment cost is at = log πθ(yt st) − log πT(yt st), and the teacher–reference log-ratio is rt = log πT(yt st) − log πref(yt st). REOPD constructs a low-rank k3 discrepancy proxy xb,i,t = log πT(yi,t si,t) − log πθ(yi,t si,t), and δ̂b,i,t = exp(xb,i,t) − xb,i,t − 1. The compatibility weight is then qb,i,t = exp(−δ̂b,i,t / τ), with τ > 0. Because δ̂b,i,t ≥ 0, the weight lies in (0, 1] in exact arithmetic. A small sampled discrepancy gives qb,i,t ≈ 1 and retains most of the extrapolation residual, whereas a large discrepancy yields a smaller weight. The temperature τ controls how rapidly this attenuation occurs.
The micro-batch reliable residual statistics aggregate two statistics over valid response tokens. The compatibility-weighted residual proportion is ρb = (Σ rb,i,t qb,i,t) / (Σ mi,t rb,i,t + ε), and the reliable residual scale is sb = (Σ mi,t (qb,i,t rb,i,t) squared / (Σ mi,t + ε))(1/2). REOPD maintains exponential moving averages: z̄b = β z̄b−1 + (1 − β) zb, for z ∈ ρ, s. The bounded micro-batch extrapolation budget is γ̃b = clip(B0 ρ̄b / (s̄b + ε), 0, γmax), and the smoothed budget is γb = βγ γb−1 + (1 − βγ) γ̃b. The effective extrapolation coefficient is λb,i,t = 1 + γb qb,i,t, and the token cost is Cb,i,t = ab,i,t − γb qb,i,t rb,i,t. The PPO-style actor update uses the negative token cost as its advantage: AREOPD b,i,t = −Cb,i,t. In exact arithmetic, 1 ≤ λb,i,t ≤ 1 + γmax.
The paper evaluates REOPD on mathematical reasoning, code generation, and mixed-domain multi-teacher distillation. The student and reference policy are initialized from Qwen3-4B, with task-specialized Qwen3-4B non-thinking policies as teachers. The mathematics run uses 57,046 level-6 examples from the filtered DeepMath-103K training set, while the code run uses 25,276 examples from the Eurus code training split. The multi-teacher set contains 25,276 examples from each domain. Baselines include standard OPD (λ = 1) and fixed-coefficient ExOPD at λ = 1.25.
For mathematics, evaluation uses AIME 2024, AIME 2025, HMMT February 2025, and HMMT November 2025, with 32 responses per problem and pooled sample accuracy over 3,840 completions. For code, evaluation uses HumanEval+, MBPP+, and LiveCodeBench v6 test6, with four responses per problem and pass@1, with aggregate code score as task-count-weighted accuracy over 2,868 completions.
Main results show that in single-teacher distillation, REOPD reaches 47.66% pooled sample accuracy on mathematics, exceeding both OPD at 46.28% and ExOPD at λ = 1.25 at 47.47%. On code generation, REOPD obtains 63.45% weighted accuracy and remains comparable to G-OPD; under the common λ = 1.25 baseline, it exceeds ExOPD at 61.72% and OPD at 62.55%. In multi-teacher distillation, REOPD reaches 47.01% mathematics sample accuracy and 63.32% code weighted accuracy, both exceeding OPD at 46.43% and 61.99%, as well as ExOPD at λ = 1.25 at 46.98% and 62.90%.
Sensitivity analysis shows that the best fixed coefficient varies across settings: single-teacher mathematics peaks at λ = 1.25 with 47.47%, single-teacher code peaks at λ = 1.5 with 63.60%, multi-teacher mathematics ties at λ = 1.25 and 1.75 at 46.98%, and multi-teacher code is best at λ = 1.25 at 62.90%. REOPD exceeds the corresponding best fixed-coefficient result by +0.19, −0.15, +0.03, and +0.42 percentage points on single-teacher mathematics, single-teacher code, multi-teacher mathematics, and multi-teacher code, respectively.
Ablation studies in single-teacher mathematics show that removing token compatibility (no q) reduces pooled accuracy from 47.66% to 43.39%, a drop of 4.27 percentage points. The no bound variant obtains 47.16%, 0.50 percentage points below Full REOPD. The no batch sweep with fixed λ0 shows accuracy ranging from 46.48% to 47.66%, with the best at λ0 = 1.25 matching Full REOPD. The ablation identifies token-level compatibility as the key component, the explicit bound as a modest safeguard, and the online micro-batch budget as providing adaptive control without fixed-coefficient selection.
Controller dynamics show that REOPD does not use a constant effective coefficient. Over the first ten steps, the mean budget γ is 0.608, 0.953, and 0.957 for mathematics, code, and multi-teacher training, respectively; over the final ten steps, these increase to 1.000, 0.986, and 0.990. Mean token compatibility increases from 0.774/0.745/0.754 to 0.850/0.827/0.834, and the mean effective coefficient increases from 1.477/1.710/1.720 to 1.850/1.815/1.826.
Limitations include that the compatibility weight is a sampled student–teacher discrepancy proxy rather than a correctness estimator; it cannot identify trajectories on which the student and teacher agree but are both wrong. REOPD retains controller choices such as τ, γmax, and the calibration of B0; its statistics depend on micro-batch composition, and the budget approaches its upper bound late in training.
Improvements for AI systems
Based on the paper, here are the specific improvements I can implement in AI systems:
1. Adaptive Reward Extrapolation Control
-
Replace fixed global scalar λ (as in ExOPD) with a dynamic, per-token effective coefficient λ b,i,t = 1 + γ b · q b,i,t
-
This prevents overfitting to extreme teacher–reference log-ratio peaks, reducing reward hacking and training instability
-
The system automatically adjusts extrapolation strength per batch based on reliability statistics, eliminating costly domain-specific λ sweeps
2. Token-Level Compatibility Weighting
-
Use q b,i,t = exp(−δ̂ b,i,t / τ) to attenuate extrapolation on tokens where student–teacher discrepancy is high
-
This ensures the student only amplifies residuals it can reliably learn from, avoiding aggressive fitting on noisy or unreliable tokens
-
The weight naturally decays to near-zero for large discrepancies, providing a safety mechanism against pathological updates
3. Bounded Micro-Batch Budget with EMA Smoothing
-
Implement γ̃ b = clip(B0·ρ̄ b / (s̄ b + ε), 0, γ max) to cap extrapolation strength per batch
-
Exponential moving averages (β = 0.9) smooth budget and statistics, preventing abrupt jumps and stabilizing training
-
The bound γ max (e.g., 1.0) guarantees the effective coefficient stays in [1, 2], ensuring updates remain conservative
4. Zero-Overhead Integration
-
Reuse existing student, teacher, and reference log-probabilities from G-OPD—no verifier, reward model, value model, or extra rollouts needed
-
This makes the improvement directly applicable to any existing on-policy distillation pipeline with minimal code changes
5. Multi-Teacher Robustness
-
The adaptive mechanism automatically handles mixed-domain distillation (e.g., math + code) without manual per-domain λ tuning
-
It outperforms fixed-coefficient baselines across all tested domains, including cases where the optimal λ differs (e.g., math best at 1.25, code best at 1.5)
Improved AI System Capabilities:
-
Stable post-training: Achieves higher accuracy (e.g., +0.19% on math, +0.42% on multi-teacher code) than the best fixed-coefficient ExOPD, with no manual hyperparameter search
-
Self-regulating: Automatically increases extrapolation strength as training progresses (mean γ rises from 0.6 to 1.0), adapting to the student’s growing reliability
-
Robust to domain shifts: Works across math reasoning, code generation, and mixed-domain settings without re-tuning
-
Safer distillation: Prevents catastrophic forgetting or reward hacking by capping updates and down-weighting unreliable tokens, making it suitable for production deployment
Sources
- Prune-OPD: Efficient and Reliable On-Policy Distillation for Long-Horizon Reasoning
- Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
- SCOPE: Signal-Calibrated On-Policy Distillation Enhancement with Dual-Path Adaptive Weighting
- TIP: Token Importance in On-Policy Distillation
- SG-OPD: Sign-Gated On-Policy Distillation via Sign-Consistency Gating and Phased Teacher Sampling
- Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation
- Reward-Gated On-Policy Distillation
- Distilling the Knowledge in a Neural Network
- Proximal Policy Optimization Algorithms
- Qwen3 Technical Report
- DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
- Advancing LLM Reasoning Generalists with Preference Trees
- Evaluating Large Language Models Trained on Code
- Program Synthesis with Large Language Models
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation
- LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks