Learning to Stop without Learning to Stop: Self-Supervised Confidence Training Improves Reasoning Efficiency
cs.AI, cs.CL, cs.LG
Submitted: 2026-09-25
Updated: 2026-09-25
Terminology
Sources
- L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
- Training Language Models to Reason Efficiently
- Llama-Nemotron: Efficient Reasoning Models
- Confidence-Aware Alignment Makes Reasoning LLMs More Reliable
- Training Verifiers to Solve Math Word Problems
- Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty
- Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs
- Deep Think with Confidence
- Understanding and Mitigating Premature Confidence for Better LLM Reasoning
- Your Models Have Thought Enough: Training Large Reasoning Models to Stop Overthinking
- DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
- ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning
- Does Your Reasoning Model Implicitly Know When to Stop Thinking?
- LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
- Verbalized Confidence Triggers Self-Verification: Emergent Behavior Without Explicit Reasoning Supervision
- Understanding the Thinking Process of Reasoning Models: A Perspective from Schoenfeld's Episode Theory
- Reinforcement Learning with Metacognitive Feedback Elicits Faithful Uncertainty Expression in LLMs
- DLER: Doing Length pEnalty Right - Incentivizing More Intelligence per Token via Reinforcement Learning
- Answer Convergence as a Signal for Early Stopping in Reasoning
- Early Stopping Chain-of-thoughts in Large Language Models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection