When Confidence Signals Disagree: Local and Global Confidence in Autoregressive Language Models
cs.LG, cs.AI
Submitted: 2026-09-15
Updated: 2026-09-15
Code: https://github.com/julioadl/stochastic-parrots
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: Modern predictive systems expose multiple quantities that are commonly interpreted as measures of confidence.
Terminology
Abstract
Modern predictive systems expose multiple quantities that are commonly interpreted as measures of confidence. However, these quantities can summarize different aspects of the predictive process. This distinction matters when confidence is used to evaluate reliability or inform downstream oversight and control. We investigate whether different confidence readouts are empirically interchangeable in an autoregressive language model by comparing local confidence, defined from the probability of the greedy-selected answer token, with global confidence, defined from modal-answer frequency under repeated sampling. Across MMLU and ARC Challenge, the two signals are weakly correlated and differ substantially in their association with correctness: global confidence is moderately associated with correctness, whereas local confidence shows little association. We further test whether question-level disagreement between the signals is associated with sampling instability. On ARC, larger local--global confidence gaps are associated with higher answer entropy, more distinct sampled answers, and lower modal-answer concentration. The gap--entropy association persists when disagreement and instability are estimated from disjoint stochastic samples, indicating that it is not explained by shared finite-sample variation. The corresponding relationship is substantially weaker on MMLU, where only 4% of questions exhibit sampling instability. These results show that confidence readouts derived from the same predictive system are not empirically interchangeable and that their disagreement can provide a diagnostic of unstable sampling behavior. Confidence should therefore be treated as an explicitly defined measurement rather than as a single intrinsic scalar property of a model, particularly when it is used to inform downstream evaluation, oversight, or control.
Sources
- Discovering Latent Knowledge in Language Models Without Supervision
- Universal Self-Consistency for Large Language Model Generation
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- The Llama 3 Herd of Models
- On Calibration of Modern Neural Networks
- Measuring Massive Multitask Language Understanding
- How Can We Know When Language Models Know? On the Calibration of Language Models for Question Answering
- Robots That Ask For Help: Uncertainty Alignment for Large Language Model Planners
- Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
- Language Models (Mostly) Know What They Know
- Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation
- Confidence Under the Hood: An Investigation into the Confidence-Probability Alignment in Large Language Models
- Self-Consistency Improves Chain of Thought Reasoning in Language Models
- Teaching Models to Express Their Uncertainty in Words
- Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs
- Uncertainty Estimation in Autoregressive Structured Prediction
- The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets
- Calibrating Sequence likelihood Improves Conditional Language Generation
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks