A Recipe for Long-Context Reasoning in Large Language Models via On-Policy Optimization and Distillation

arXiv:2605.12227 · cs.CL · Submitted 2026-05-12 · Read on arXiv

cs.CL

Submitted: 2026-05-12

Updated: 2026-09-10

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Terminology

Sources

Related papers