Rethinking Training-Inference Mismatch in LLM Reinforcement Learning: Where It Arises and How to Correct It

arXiv:2609.32444 · cs.LG, cs.AI · Submitted 2026-09-26 · Read on arXiv

cs.LG, cs.AI

Submitted: 2026-09-26

Updated: 2026-09-26

Code: https://github.com/kzhao5/CIS-RL

Terminology

Sources

Related papers