On-Policy Delta Distillation for Multilingual Math Reasoning
Byeongho Heo, Jaehui Hwang, Sangdoo Yun, Dongyoon Han
cs.CL, cs.LG
Submitted: 2026-08-06
Comments: 9 pages, 3 figures, 10 tables
Code: https://github.com/naver-ai/opd2
License: http://creativecommons.org/licenses/by/4.0/
The gist: On-Policy Distillation (OPD) is emerging as a promising alternative to reinforcement learning for LLM post-training, yet its effectiveness in multilingual settings remains underexplored.
Terminology
Abstract
On-Policy Distillation (OPD) is emerging as a promising alternative to reinforcement learning for LLM post-training, yet its effectiveness in multilingual settings remains underexplored. We study OPD and its advanced variant, On-Policy Delta Distillation (OPD squared), for mathematical reasoning in English, Korean, and Japanese. OPD squared improves OPD by using the probability gap between a post-trained teacher and its base model as the learning signal. Experiments with Qwen3 show that OPD squared consistently outperforms the original OPD, with particularly strong improvements in Korean and Japanese, and generally narrows the English-Korean performance gap. We further find that English-only OPD can also increase performance for Korean and Japanese, but often shifts the responses toward English, highlighting the importance of multilingual data to preserving target-language responses.
Sources
- Qwen3 Technical Report
- MiMo-V2-Flash Technical Report
- Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation
- On-Policy Delta Distillation
- A Survey of On-Policy Distillation for Large Language Models
- Crosslingual On-Policy Self-Distillation for Multilingual Reasoning
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Nemotron-Math: Efficient Long-Context Distillation of Mathematical Reasoning from Multi-Mode Supervision
- Language Models are Multilingual Chain-of-Thought Reasoners
- BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering