Post-Training Large Language Models via Reinforcement Learning from Self-Feedback
cs.CL, cs.AI
Submitted: 2025-07-29
Updated: 2026-09-15
Code: https://github.com/huggingface/trl
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large Language Models (LLMs) often produce plausible but poorly-calibrated answers, limiting their reliability on reasoning-intensive tasks.
Terminology
Abstract
Large Language Models (LLMs) often produce plausible but poorly-calibrated answers, limiting their reliability on reasoning-intensive tasks. Recent research suggests that Chain-of-Thought (CoT) reasoning paths are inherent in pre-trained LLMs and can be elicited by simply altering the decoding process, where the presence of a CoT path correlates with higher answer confidence. Building on these insights, we present Reinforcement Learning from Self-Feedback (RLSF), a post-training stage that utilises the model's intrinsic confidence as a self-generated reward. By generating multiple CoT decoding beams from a frozen LLM, we compute the confidence of each final answer span and rank the resulting traces accordingly to create synthetic preferences. These preferences are subsequently utilised to fine-tune the policy through standard preference optimisation, requiring no human labels, gold answers, or externally curated rewards. RLSF simultaneously (i) refines the model's probability estimates--restoring well-behaved calibration--and (ii) strengthens step-by-step reasoning, yielding improved performance on arithmetic reasoning and multiple-choice question answering. By converting a model's own uncertainty into structured self-feedback, RLSF affirms reinforcement learning on intrinsic model behaviour as a principled and data-efficient component of the LLM post-training pipeline. Our results demonstrate that leveraging these inherent reasoning capabilities provides a robust path for enhancing model reliability without manual prompt engineering or external supervision.
Sources
- Constitutional AI: Harmlessness from AI Feedback
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Training Verifiers to Solve Math Word Problems
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Quantile Regression for Distributional Reward Models in RLHF
- Beyond Accuracy: The Role of Calibration in Self-Improving Large Language Models
- LLM Post-Training: A Deep Dive into Reasoning Large Language Models
- Tulu 3: Pushing Frontiers in Open Language Model Post-Training
- RewardBench: Evaluating Reward Models for Language Modeling
- Uncertainty-aware Reward Model: Teaching Reward Models to Know What is Unknown
- GPT-4 Technical Report
- OpenAI o1 System Card
- Gemma 2: Improving Open Language Models at a Practical Size
- Chain-of-Thought Reasoning Without Prompting
- Proximal Policy Optimization Algorithms
- Rewarding Doubt: A Reinforcement Learning Approach to Calibrated Confidence Expression of Large Language Models
- Inverse-RLignment: Large Language Model Alignment from Demonstrations through Inverse Reinforcement Learning
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering