Scaling Properties of Same-Family On-Policy Distillation
cs.LG
Submitted: 2026-09-26
Updated: 2026-09-26
Code: https://github.com/chrisliu298/awesome-on-policy-distillation
Terminology
Sources
- The Mechanism of Weak-to-Strong Generalization: Feature Elicitation from Latent Knowledge
- Unified Neural Scaling Laws
- H$^2$SD: Hybrid Hindsight Self-Distillation
- Training Verifiers to Solve Math Word Problems
- Weak-to-Strong Generalization via Direct On-Policy Distillation
- Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes
- Measuring Mathematical Problem Solving With the MATH Dataset
- Distilling the Knowledge in a Neural Network
- Training Compute-Optimal Large Language Models
- Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention
- High-dimensional Analysis of Knowledge Distillation: Weak-to-Strong Generalization and Scaling Laws
- Scaling Laws for Neural Language Models
- Scaling Reasoning Efficiently via Relaxed On-Policy Distillation
- Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
- On-policy Distillation with Verifiable Reward
- Strong Teacher Not Needed? On Distillation in LLM Pretraining
- Olmo 3
- Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation
- Qwen2.5 Technical Report
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks