Calibration, Not Answer Selection: Distilling Internal Confidence in Reasoning Models
cs.CL
Submitted: 2026-09-27
Updated: 2026-09-27
Code: https://github.com/yale-nlp/RLMF
Terminology
Sources
- Understanding intermediate layers using linear classifier probes
- The Internal State of an LLM Knows When It's Lying
- Perceptions of Linguistic Uncertainty by Language Models and Humans
- Discovering Latent Knowledge in Language Models Without Supervision
- Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Trace Length is a Simple Uncertainty Signal in Reasoning Models
- High Accuracy, Less Talk (HALT): Reliable LLMs through Capability-Aligned Finetuning
- Inside-Out: Hidden Factual Knowledge in LLMs
- Thinking to Recall: How Reasoning Unlocks Parametric Knowledge in LLMs
- GLM-5: from Vibe Coding to Agentic Engineering
- LLMs Should Express Uncertainty Explicitly
- SimpleQA Verified: A Reliable Factuality Benchmark to Measure Parametric Knowledge
- Distilling the Knowledge in a Neural Network
- Calibrating Verbal Uncertainty as a Linear Feature to Reduce Hallucinations
- TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension
- Language Models (Mostly) Know What They Know
- Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation
- Reported Confidence in LLMs Tracks Commitment More Than Correctness
- Calibrating Overconfidence Without Sacrificing Confidence: Probe-Conditioned Head Intervention for LLMs
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering