How do LLMs Compute Verbal Confidence
cs.CL, cs.AI, cs.LG
Submitted: 2026-03-18
Updated: 2026-09-18
License: http://creativecommons.org/licenses/by/4.0/
The gist: Verbal confidence -- prompting LLMs to state their confidence as a number or category -- is widely used to extract uncertainty estimates from black-box models.
Terminology
Abstract
Verbal confidence -- prompting LLMs to state their confidence as a number or category -- is widely used to extract uncertainty estimates from black-box models. However, how LLMs internally generate such scores remains unknown. We address two questions: first, when confidence is computed -- just-in-time when requested, or automatically during answer generation and cached for later retrieval; and second, what verbal confidence represents -- token log-probabilities, or a richer evaluation of answer quality? Focusing on Gemma 3 27B (across TriviaQA, BigMath, and MMLU), Qwen 2.5 7B, and the reasoning model Magistral Small 24B, we provide convergent evidence for cached retrieval. Activation steering, patching, noising, and swap experiments reveal that confidence representations emerge at answer-adjacent positions before appearing at the verbalization site. Attention blocking pinpoints the information flow: confidence is gathered from answer tokens, cached at the first post-answer position, then retrieved for output. Critically, linear probing and variance partitioning reveal that these cached representations explain substantial variance in verbal confidence beyond token log-probabilities, suggesting a richer answer-quality evaluation rather than a simple fluency readout. These findings demonstrate that verbal confidence reflects automatic, sophisticated self-evaluation -- not post-hoc reconstruction -- with implications for understanding metacognition in LLMs and improving calibration.
Sources
- Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models
- How to use and interpret activation patching
- Steering Evaluation-Aware Language Models to Act Like They Are Deployed
- The Internal State of an LLM Knows When It's Lying
- TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension
- Language Models (Mostly) Know What They Know
- Discovering Latent Knowledge in Language Models Without Supervision
- Trace Length is a Simple Uncertainty Signal in Reasoning Models
- How LLMs Detect and Correct Their Own Errors: The Role of Internal Confidence Signals
- A Survey of Confidence Estimation and Calibration in Large Language Models
- Dissecting Recall of Factual Associations in Auto-Regressive Language Models
- LLM Microscope: What Model Internals Reveal About Answer Correctness and Context Utilization
- Reasoning about Uncertainty: Do Reasoning Models Know When They Don't Know?
- Steering Llama 2 via Contrastive Activation Addition
- Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback
- Steering Language Models With Activation Engineering
- A Practical Review of Mechanistic Interpretability for Transformer-Based Language Models
- A Stitch in Time Saves Nine: Detecting and Mitigating Hallucinations of LLMs by Validating Low-Confidence Generation
- Base Models Know How to Reason, Thinking Models Learn When
- Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering