Rethinking Training-Inference Mismatch in LLM Reinforcement Learning: Where It Arises and How to Correct It
cs.LG, cs.AI
Submitted: 2026-09-26
Updated: 2026-09-26
Code: https://github.com/kzhao5/CIS-RL
Terminology
Sources
- Program Synthesis with Large Language Models
- MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
- Evaluating Large Language Models Trained on Code
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
- Measuring Massive Multitask Language Understanding
- Measuring Mathematical Problem Solving With the MATH Dataset
- Solving Quantitative Reasoning Problems with Language Models
- Trust Region Masking for Long-Horizon LLM Reinforcement Learning
- Ling and Ring 2.6 Technical Report: Efficient and Instant Agentic Intelligence at Trillion-Parameter Scale
- Every Step Evolves: Scaling Reinforcement Learning for Trillion-Scale Thinking Model
- Stabilizing MoE Reinforcement Learning by Aligning Training and Inference Routers
- Defeating the Training-Inference Mismatch via FP16
- Qwen3 Technical Report
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Beyond Precision: Training-Inference Mismatch is an Optimization Problem and Simple LR Scheduling Fixes It
- Stabilizing Reinforcement Learning with LLMs: Formulation and Practices
- Group Sequence Policy Optimization
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks