When Words Fall Short: Iterative Synergy Between Verbalized Reasoning and Hidden Features for LLM Confidence Estimation
cs.CL
Submitted: 2026-09-28
Updated: 2026-09-28
Code: https://github.com/xyk829/ipoet
Terminology
Sources
- Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models
- Rewarding Doubt: A Reinforcement Learning Approach to Calibrated Confidence Expression of Large Language Models
- Uncalibrated Reasoning: GRPO Induces Overconfidence for Stochastic Outcomes
- INSIDE: LLMs' Internal States Retain the Power of Hallucination Detection
- Training Verifiers to Solve Math Word Problems
- Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty
- The Llama 3 Herd of Models
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- LLMs Should Express Uncertainty Explicitly
- Language Model Cascades: Token-level uncertainty and beyond
- Measuring Mathematical Problem Solving With the MATH Dataset
- OpenAI o1 System Card
- Language Models (Mostly) Know What They Know
- How do LLMs Compute Verbal Confidence
- Tulu 3: Pushing Frontiers in Open Language Model Post-Training
- ORCE: Order-Aware Alignment of Verbalized Confidence in Large Language Models
- Agent Hospital: A Simulacrum of Hospital with Evolvable Medical Agents
- Decoupling Reasoning and Confidence: Resurrecting Calibration in Reinforcement Learning from Verifiable Rewards
- Annotation-Efficient Universal Honesty Alignment
- GPQA: A Graduate-Level Google-Proof Q&A Benchmark
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering