SIPO: Unifying Reinforcement Learning with On-Policy Self-Distillation

arXiv:2609.36742 · cs.AI, cs.CL · Submitted 2026-09-29 · Read on arXiv

cs.AI, cs.CL

Submitted: 2026-09-29

Updated: 2026-09-29

Code: https://github.com/Yueeeeeeee/SIPO

Terminology

Sources

Related papers