Future Policy Approximation for Offline Reinforcement Learning Improves Mathematical Reasoning

arXiv:2509.19893 · cs.CL · Submitted 2026-08-19 · Read on arXiv

Minjae Oh, Yunho Choi, Dongmin Choi, Yohan Jo

cs.CL

Submitted: 2026-08-19

Updated: 2026-08-20

Code: https://github.com/huggingface/trl

Terminology

Sources

Related papers