ReST-RL: Reinforcing LLM Reasoning through Unified Self-Training and Value-Guided Search
cs.AI, cs.LG
Submitted: 2025-08-27
Updated: 2026-09-07
Comments: 29 pages, 5 figures
Code: https://github.com/THUDM/ReST-RL
License: http://creativecommons.org/licenses/by/4.0/
The gist: With respect to improving the reasoning accuracy of LLMs, the representative reinforcement learning (RL) method - Group Relative Policy Optimization (GRPO) - has achieved critical success, yet it
Terminology
Abstract
With respect to improving the reasoning accuracy of LLMs, the representative reinforcement learning (RL) method - Group Relative Policy Optimization (GRPO) - has achieved critical success, yet it still suffers from the issue of insignificant reward signals. This paper introduces ReST-RL, a unified Reinforced Self-Training (ReST) policy-value framework that reconnects policy optimization and value-guided search to improve LLM reasoning ability. Firstly, ReST-GRPO adopts an optimized ReST-style algorithm to reshape the policy-induced trajectory distribution by increasing the reward variance of GRPO sampling and exposing the policy to more informative partial states, thereby improving training efficiency and effectiveness. Then, we further introduce a decoding optimization method, VM-MCTS, which trains a Value Model (VM) from self-collected Monte-Carlo Tree Search (MCTS) targets and deploys it through an adapted MCTS algorithm to provide precise process signals and verification scores, further enhancing reasoning accuracy. These two stages are internally dependent - ReST-GRPO yields higher-quality trajectories for value learning with VM-MCTS, which in turn enables more effective inference-time search. We validate our framework on multiple coding benchmarks (e.g., APPS, BigCodeBench, and HumanEval), where it significantly outperforms other reinforcement training baselines (naive GRPO, DAPO, and ReST-DPO), as well as decoding and verification baselines (e.g., PRM-BoN and ORM-MCTS), indicating its power to strengthen LLM reasoning capability. Moreover, we further evaluate ReST-RL on out-of-domain math and science reasoning tasks, where it achieves improved performance without target-domain tuning and favorable end-to-end efficiency trade-offs, providing preliminary transfer evidence beyond our primary coding domain.
Sources
- Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models
- OpenCodeInterpreter: Integrating Code Generation with Execution and Refinement
- Qwen2.5-Coder Technical Report
- Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection