Reinforcing Step-level Reasoning for Effective Self-Correction in LLMs
Nanyang Technological University · National University of Singapore · VinUniversity · Institute for Infocomm Research
cs.CL, cs.AI
Submitted: 2026-08-12
Updated: 2026-08-31
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 87/100
The gist: The paper introduces Self-Fix Step-DPO (SFS-DPO), a two-stage reinforcement learning framework designed to improve self-correction capabilities in large language models (LLMs) for mathematical
Terminology
Summary
The paper introduces Self-Fix Step-DPO (SFS-DPO), a two-stage reinforcement learning framework designed to improve self-correction capabilities in large language models (LLMs) for mathematical reasoning. The authors argue that existing step-level training methods primarily optimize preferences over better reasoning continuations without explicitly correcting erroneous steps already generated. To address this, they propose a framework where the first stage strengthens step-level reasoning via step-level preference optimization, and the second stage explicitly trains models to self-verify and self-correct.
The paper states: "We propose Self-Fix Step-DPO (SFS-DPO), a reinforcement learning based, two-stage framework for step-level self-verification and self-correction. The first stage strengthens step-level reasoning via step-level preference optimization, while the second stage explicitly trains models to self-verify and self-correct."
The authors also introduce a teacher-assisted variant, SFS-DPO-R, which incorporates explanatory rationales for error verification to provide stronger corrective signals.
In this variant, a strong teacher model (GPT-4o) generates explicit explanations of why a previous step is incorrect, and these rationales are inserted between the self-correction signal and the correct reasoning step.
The methodology is described as follows: "We formulate self-correction by allowing each step sj to take one of three types, SOLUTION STEP, ERROR-DETECTION STEP, FIXED STEP, corresponding respectively to a standard reasoning step, an explicit signal (optionally with reasoning) indicating that the previous step is incorrect, or a corrected version of an erroneous step."
The initialization stage uses a reinforcement learning objective: "In the first stage, we employ a reinforcement learning objective to initialize the model... Specifically, it aims to maximize the likelihood of a correct SOLUTION STEP s+k while minimizing that of the incorrect SOLUTION STEP s−k. The second stage constructs a preference objective that
favors a self-corrected continuation c+k over the subsequent incorrect step s−k+1 that would be generated if the error remains unaddressed."
Experiments are conducted on seven LLMs, including DeepSeekMath-7B, Qwen2-7B, Qwen2-7B-Instruct, Qwen2.5-Math-7B-Instruct, Qwen3-8B, Llama-3.1-8B-Instruct, and Qwen2.5-14B-Instruct. The evaluation covers in-domain benchmarks (MATH, GSM8K) and out-of-domain benchmarks (GK2023, OCW). The results show that SFS-DPO and SFS-DPO-R consistently outperform prior step-level training baselines.
Specifically, the paper reports: Averaged across seven backbones, SFS-DPO improves accuracy by 1.11% on MATH and 0.69% on GSM8K, while SFS-DPO-R further increases the gains to 1.36% and 0.87%, respectively.
For out-of-domain results, Qwen2-7B-Instruct achieves a substantial 6.3% improvement on GK2023 under SFS-DPO, while Qwen2.5-Math-7B-Instruct achieves the largest overall OOD gains, reaching 10.9% on GK2023 and 8.8% on OCW.
The paper also analyzes the role of the initialization stage, finding that Removing the initialization stage leads to clear performance degradation, showing that directly optimizing self-correction signals without a strong reasoning foundation is insufficient.
Additionally, joint training is also detrimental, suggesting that step-wise reasoning and self-correction are better learned sequentially rather than as a single mixed objective.
Regarding self-correction rate, the authors note: SC rate alone is insufficient as a standalone indicator of self-correction quality. What ultimately matters is whether self-correction behavior is positively aligned with task performance.
They observe that a high SC rate does not necessarily translate into better accuracy,
and that their methods learn to apply self-correction more selectively and effectively.
The paper's contributions are summarized as: (1) proposing an RL-based, step-level self-correction training framework with two variants, SFS-DPO and SFS-DPO-R; (2) conducting comprehensive in-domain and out-domain experiments demonstrating consistent improvements over baselines across multiple LLMs; and (3) analyzing self-correction behavior in the framework, providing empirical insights into its frequency, effectiveness, and relationship with reasoning performance.
The authors conclude: "These results highlight that explicitly modeling how errors are identified and fixed, rather than merely preferring better continuations, is critical for robust math reasoning, positioning step-wise self-correction as a promising direction for improving the reliability of large language models."
Improvements for AI systems
Improvements to AI Systems:
-
Explicit Error-Detection and Correction Mechanism: Integrate a three-step token type system (SOLUTION STEP, ERROR-DETECTION STEP, FIXED STEP) into the model's decoding process. This allows the AI to explicitly flag its own mistakes mid-reasoning, generate a corrective signal, and then produce a revised step—rather than silently continuing from an error. The improved system can halt its own flawed reasoning chains and visibly correct them, improving final answer accuracy on multi-step math problems.
-
Two-Stage Sequential Training Pipeline: Implement the two-stage training approach where the model first learns robust step-level reasoning via preference optimization (stage 1), and only then learns self-verification and correction (stage 2). This prevents the model from attempting self-correction without a strong reasoning foundation. The improved system will exhibit more stable learning and higher final accuracy than if trained on a mixed objective, as joint training was shown to be detrimental.
-
Teacher-Assisted Rationale Injection (SFS-DPO-R): Use a stronger teacher model (e.g., GPT-4o) to generate explanatory rationales for why a previous step is incorrect, and insert these rationales between the error-detection signal and the corrected step during training. The improved AI system will learn not just that a step is wrong, but why it is wrong, enabling it to generalize error-correction to unseen problem types and out-of-domain benchmarks (e.g., achieving up to 10.9% improvement on GK2023).
-
Selective Self-Correction Policy: Train the model to apply self-correction only when it is beneficial, rather than frequently. The framework's preference objective (favoring corrected continuations over uncorrected erroneous steps) teaches the model to discriminate between errors that are worth fixing and those that are not. The improved system will have a lower but more effective self-correction rate, leading to higher task performance without overcorrecting or introducing new errors.
-
Sequential Error-Repair Optimization: Replace the standard step-level preference optimization (which only prefers better continuations) with a preference that explicitly compares a self-corrected continuation against the continuation that would result from ignoring the error. The improved AI system will learn to prioritize repairing the immediate erroneous step, reducing error propagation across long reasoning chains—a key factor for robust performance on complex math problems.
-
Backbone-Agnostic Fine-Tuning Recipe: Apply the SFS-DPO framework as a drop-in fine-tuning method across diverse model families (DeepSeekMath, Qwen, Llama) without architectural changes. The improved system can be deployed on any existing LLM to boost its mathematical reasoning reliability, with consistent gains (e.g., +1.36% on MATH, +0.87% on GSM8K averaged across seven backbones) and no need for custom model designs.
-
Out-of-Domain Generalization for Reasoning: Leverage the teacher rationales and sequential training to improve the model's ability to detect and fix errors in unfamiliar problem settings. The improved system will show significant accuracy gains on out-of-domain benchmarks (e.g., Qwen2.5-Math-7B-Instruct achieving +10.9% on GK2023 and +8.8% on OCW), demonstrating that explicit self-correction training transfers beyond the training distribution.
Abstract
Achieving effective self-correction, where models verify and correct their own mistakes, remains a fundamental challenge for large language models (LLMs). In this work, we propose Self-Fix Step-DPO (SFS-DPO), a reinforcement learning based, two-stage framework for step-level self-verification and self-correction. The first stage strengthens step-level reasoning via step-level preference optimization, while the second stage explicitly trains models to self-verify and self-correct. We further introduce a teacher-assisted variant, SFS-DPO-R, which incorporates explanatory rationales for error verification to provide stronger corrective signals. Comprehensive in-domain and out-of-domain evaluations across multiple LLMs demonstrate that SFS-DPO and SFS-DPO-R consistently outperform prior step-level training baselines. Our analysis further reveals improvements in self-correction frequency and effectiveness, highlighting the importance of strengthening step-level reasoning for robust performance.
Sources
- ORPO: Monolithic Preference Optimization without Reference Model
- GPT-4o System Card
- ReCUT: Balancing Reasoning Length and Accuracy in LLMs via Stepwise Trails and Preference Optimization
- Training Language Models to Self-Correct via Reinforcement Learning
- Qwen Technical Report
- Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
- Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models
- Training Verifiers to Solve Math Word Problems
- The Llama 3 Herd of Models
- Qwen2.5 Technical Report
- S$^2$R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning
- Evolving LLMs' Self-Refinement Capability via Synergistic Training-Inference Optimization
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Boosting LLM Reasoning via Spontaneous Self-Correction
- Qwen3 Technical Report
- Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering