What are Key Factors for Updates in RL for LLM Reasoning?
cs.CL
Submitted: 2026-06-21
Updated: 2026-09-19
Code: https://github.com/Control-derek/ACPO
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Asymmetric REINFORCE for off-Policy Reinforcement Learning: Balancing positive and negative rewards
- MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
- SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization
- AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
- Reasoning with Exploration: An Entropy Perspective
- GPG: A Simple and Strong Reinforcement Learning Baseline for Model Reasoning
- The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
- SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning
- One-shot Entropy Minimization
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Segment Policy Optimization: Effective Segment-Level Credit Assignment in RL for Large Language Models
- OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
- Measuring Mathematical Problem Solving With the MATH Dataset
- Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model
- OpenAI o1 System Card
- Tulu 3: Pushing Frontiers in Open Language Model Post-Training
- Part I: Tricks or Traps? A Deep Dive into RL for LLM Reasoning
- Learning What Reinforcement Learning Can't: Interleaved Online Fine-Tuning for Hardest Questions
- Tapered Off-Policy REINFORCE: Stable and efficient reinforcement learning for LLMs
- Proximal Policy Optimization Algorithms
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering