On-Policy Delta Distillation for Multilingual Math Reasoning

arXiv:2608.05802 · cs.CL, cs.LG · Submitted 2026-08-06 · Read on arXiv

Byeongho Heo, Jaehui Hwang, Sangdoo Yun, Dongyoon Han

cs.CL, cs.LG

Submitted: 2026-08-06

Comments: 9 pages, 3 figures, 10 tables

Code: https://github.com/naver-ai/opd2

License: http://creativecommons.org/licenses/by/4.0/

The gist: On-Policy Distillation (OPD) is emerging as a promising alternative to reinforcement learning for LLM post-training, yet its effectiveness in multilingual settings remains underexplored.

Terminology

Abstract

On-Policy Distillation (OPD) is emerging as a promising alternative to reinforcement learning for LLM post-training, yet its effectiveness in multilingual settings remains underexplored. We study OPD and its advanced variant, On-Policy Delta Distillation (OPD squared), for mathematical reasoning in English, Korean, and Japanese. OPD squared improves OPD by using the probability gap between a post-trained teacher and its base model as the learning signal. Experiments with Qwen3 show that OPD squared consistently outperforms the original OPD, with particularly strong improvements in Korean and Japanese, and generally narrows the English-Korean performance gap. We further find that English-only OPD can also increase performance for Korean and Japanese, but often shifts the responses toward English, highlighting the importance of multilingual data to preserving target-language responses.

Sources

Related papers