Causal Evidence that Language Models use Confidence to Drive Behavior
cs.LG
Submitted: 2026-03-23
Updated: 2026-09-18
License: http://creativecommons.org/licenses/by/4.0/
The gist: Metacognition -- assessing the quality of one's own cognitive performance -- guides adaptive behavior across species.
Terminology
Abstract
Metacognition -- assessing the quality of one's own cognitive performance -- guides adaptive behavior across species. Substantial research demonstrates that confidence signals can be extracted from language model outputs, yet a fundamental question remains: do models actually use these signals to control behavior, such as deciding whether to answer or abstain? To investigate, we developed a four-phase paradigm. Phase 1 elicited baseline confidence estimates without an abstention option. Phase 2 revealed that LLMs apply an implicit threshold to internal confidence when deciding to abstain, with confidence effect sizes approximately an order of magnitude larger than alternative mechanisms. Phase 3 provided direct causal evidence through activation steering: boosting or suppressing confidence signals correspondingly decreased or increased abstention rates. Phase 4 extended this by systematically varying instructed thresholds, demonstrating that LLMs actively deploy confidence signals to implement abstention policies. Critically, beyond calibrated log-probability based confidence derived from the output distribution, verbal confidence independently predicted abstention across all models, despite being objectively less discriminatory of answer correctness. Activation decoding at the last pre-answer token further showed that both observable measures are lossy readouts of a richer internal representation. Together, these results suggest that abstention is not fully captured by the strength of evidence in the output distribution alone, but is better explained by the joint operation of a multidimensional internal confidence representation and threshold-based policies -- consistent with structured metacognitive control in LLMs, a capacity of growing importance as models transition to autonomous agents that must recognize their own uncertainty.
Sources
- On the Opportunities and Risks of Foundation Models
- Learning to Route LLMs with Confidence Tokens
- Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs
- A Survey of Confidence Estimation and Calibration in Large Language Models
- Steering Evaluation-Aware Language Models to Act Like They Are Deployed
- Language Models (Mostly) Know What They Know
- AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions
- How Overconfidence in Initial Choices and Underconfidence Under Criticism Modulate Change of Mind in Large Language Models
- How do LLMs Compute Verbal Confidence
- Do LLMs Know When to NOT Answer? Investigating Abstention Abilities of Large Language Models
- Steering Llama 2 via Contrastive Activation Addition
- Probabilities of Chat LLMs Are Miscalibrated but Still Predict Correctness on Multiple-Choice Q&A
- Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
- Generalization of Fine-Tuned Uncertainty Communication and Metacognition in Large Language Models
- Improving Instruction-Following in Language Models through Activation Steering
- Revisiting Uncertainty Estimation and Calibration of Large Language Models
- Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback
- Fine-Tuning Large Language Models to Appropriately Abstain with Semantic Entropy
- Uncertainty-Based Abstention in LLMs Improves Safety and Reduces Hallucinations
- Steering Language Models With Activation Engineering
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks