On Calibration of Large Language Models: From Response To Capability
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "On Calibration of Large Language Models".
Jane: Large language models are widely deployed as general-purpose problem solvers, making accurate confidence estimation critical for reliable use, and this work introduces capability calibration to address its limitations.
Tom: First, who's behind it and why it matters.
Paper summary: Tom: So we're diving into "On Calibration of Large Language Models: From Response To Capability," and what the authors are saying is that current confidence estimation methods aren't quite hitting the mark for how we actually use these models in the real world. It looks like they're proposing a shift from focusing on individual answers to understanding the model's overall ability to solve a query.
Jane: That makes sense, Tom; it sounds like they are addressing a fundamental disconnect between what we thought was useful and what actually works when we use LLMs for complex tasks. The core idea seems to be that response-level confidence doesn't capture the full picture of the model's potential accuracy on a query overall.
Lu: I think this is really interesting because it touches on how these stochastic decoding processes affect our understanding of model performance, which is something we see everywhere in complex systems at Tsinghua. The paper argues that this mismatch between response calibration and capability calibration stems directly from the way modern LLMs generate text through sampling.
Meng: From an engineering standpoint, if response-level confidence is misleading us, it impacts how we allocate resources across different queries or tasks where accuracy matters significantly. We need a metric that reflects the model's true expected success rate rather than just the luck of one specific output.
Lalam: If we think about this from a cultural perspective, capability calibration could help us build trust in AI systems by providing a more honest assessment of their actual performance potential when facing real-world problems. It moves us past just checking if a single sentence is right to assessing if the system can handle the whole task reliably.
Tom: Exactly, Lalam; that shift from checking one answer to checking overall capability is what makes this paper so important for practical deployment. The authors formally define this new capability calibration as targeting the model’s expected accuracy on a query, which they define mathematically as the probability that a response sampled from the model’s output distribution conditioned on x is correct.
Jane: To put that into simpler terms, Tom, they are moving away from asking how confident the model *is* in its single answer to asking how likely it is that the response it *will* give us will actually be right for a given question. That difference between those two notions is what they focus on.
Lu: The distinction they draw between response calibration and capability calibration is really sharp; one looks at the correctness of a single decoded output, while the other evaluates how confident the model is in answering a specific query overall. They show that these two definitions are theoretically different and also empirically different in practice.
Paper summary: Meng: That theoretical difference is what we need to address when designing systems that rely on LLMs for decision-making; if we use the wrong calibration, our downstream applications will be based on a flawed assumption about the model's reliability.
Lalam: It means that instead of just looking at the Brier score for a single response, we need to focus on estimating that expected accuracy over the entire output distribution, which is a much more robust way to gauge true utility for our users.
Tom: Right, and they quantify this difference using something called the Jensen Gap; it’s equivalent to the variance of response correctness under the model’s output distribution. That gap shows exactly how much we lose when we only look at response calibration instead of capability calibration.
Jane: So, if we look at their target for capability calibration, it's defined as µ(x, fθ), which is essentially calculating the expected accuracy of fθ given x by taking the expected value over all possible responses.
Lu: That formulation is quite powerful because it directly ties the confidence score to the model’s inherent capability concerning that specific input query x, rather than just a post-hoc guess about a single output. It forces us to consider the entire output space.
Meng: But how do we actually figure out this expected accuracy when we can't just look at it directly, which is where their empirical setup comes in; they propose estimating µ(x, fθ) through repeated sampling of "keval responses" and taking the empirical mean, setting a sample size guidance of one hundred to keep things stable.
Lalam: Sampling one hundred times per query sounds like a manageable trade-off; it allows us to get a reliable estimate without incurring an overwhelming computational cost that would make real-time applications impossible for many users.
Tom: That sampling approach seems like a practical way to approximate the true capability calibration target when we can't observe it perfectly, and they then explore several methods for producing these calibrated scores, ranging from uniform random baselines to more involved training-based probing methods.
Jane: Those different confidence estimation methods they evaluate show a lot of variation; you have things like verbalized confidence, where the model reports its own probability of being correct on the query itself, versus looking at hidden states after training on specific tasks.
Lu: It’s fascinating how they categorize these methods into training-free and training-based approaches, showing that we have a spectrum of ways to estimate this capability, from simple post-hoc guessing to methods that require more effort in terms of model probing.
Paper summary: Meng: If we consider the practical impact, the paper suggests that probing LLMs' hidden states offers a favorable tradeoff between computational cost and performance; they found that methods like "Probe" perform quite well in many in-domain settings when compared against other approaches.
Lalam: That means for deploying systems where we need good results without massive compute overhead, focusing on those model activation probes seems like a sensible path forward right now, especially since it's noted as effective for certain domains.
Tom: Indeed, the empirical results they show confirm that capability calibration leads to improved performance in simulations like pass@k prediction and inference budget allocation when compared to the baselines they tested. They even showed that using Oracle Capability-Calibrated confidence simulates groundtruth performance almost perfectly in those scenarios.
Jane: That simulation accuracy is impressive, Tom; it suggests that if we use this calibrated confidence score, our planning for tasks like pass@k success rates will be much more reliable than what we get from just using response calibration scores.
Lu: The implication here is that capability calibration provides a better signal for making high-stakes decisions about when and how to deploy these models in complex environments, moving us closer to systems that truly understand their own potential.
Meng: For the engineer on the ground, this means we can use these estimates not just for checking if one output is good, but for intelligently deciding which queries get more processing power based on their expected success rate.
Lalam: And for our culture at the startup, this research pushes us to be more rigorous and intentional about how we integrate these models into our products, demanding a higher standard of reliability from them.
Tom: So, as we wrap up this look at "On Calibration of Large Language Models: From Response To Capability," the main point is that response calibration falls short because of the randomness in LLM decoding, and capability calibration provides a more accurate measure for overall query success.
Jane: It really boils down to moving our focus from assessing a single output's correctness to understanding the model's true expected accuracy on a given input. That distinction is what they formalize so clearly in this paper.
Lu: This work lays out the theoretical groundwork for better confidence estimation, showing precisely why we need to target expected accuracy instead of just single-response correctness when dealing with these probabilistic models.
Meng: For practical deployment, the real utility lies in using that capability calibration to intelligently manage computational resources and predict performance metrics more accurately.
Lalam: It sets a new benchmark for how we should measure AI reliability, pushing us toward a system where confidence reflects true potential rather than just superficial output correctness.
Conclusion: Tom: So, we’ve just seen some heavy hitters in this paper about how we measure AI confidence, specifically focusing on moving beyond response calibration toward capability calibration.
Jane: Right, Tom; essentially the authors are showing that simply checking if a single answer is correct doesn't tell us as much about the model's overall ability to solve a query as we thought it did.
Lu: I find the mathematical framing really compelling because they formally separate response calibration, which is about an individual output, from capability calibration, which targets the probability of correctness across the entire output distribution.
Meng: From an engineering standpoint, that distinction matters because if we use response confidence for resource allocation in a system like this startup's platform, we could be making decisions based on flawed assumptions about actual performance.
Lalam: And for our culture at the startup, this research suggests a new level of rigor is needed when we assess AI reliability; it demands that our confidence metrics reflect true potential rather than just superficial output correctness.
Tom: Exactly! The authors title the piece "On Calibration of Large Language Models: From Response To Capability," and they are laying out this formal framework to show why the old way isn't sufficient for robust deployment.
Jane: It boils down to shifting our focus from assessing the correctness of one generated sentence to understanding the model’s expected accuracy on a given input query.
Lu: That shift is profound because it directly addresses the stochastic nature of LLM decoding, which is what causes this divergence between response and capability calibration targets in practice.
Meng: The implication for us practically is that when we build inference budget allocation algorithms, using these capability scores could lead to much smarter resource distribution across different user queries.
Lalam: I see a future where this level of calibration helps us build truly trustworthy AI systems because the confidence score would be a better reflection of what the AI can actually achieve on a task.
Tom: It's definitely a big deal for how we think about deploying these models, and it opens up some really interesting avenues for how we design downstream applications.
Jane: What this paper suggests is that capability calibration provides a more accurate signal for making high-stakes decisions about when and how to use these models in complex environments.
Lu: And the authors show that empirical methods like probing model hidden states offer a favorable tradeoff between computational cost and performance, which is very encouraging for real-world implementation.
Sin-Han Yang, Cheng-Kuang Wu, Chieh-Yen Lin, Yun-Nung Chen, Hung-yi Lee, Shao-Hua Sun
National Taiwan University
cs.CL, cs.AI, cs.LG
Submitted: 2026-02-14
Updated: 2026-09-29
Code: https://github.com/appier-research/llm-calibration
Importance score: 92/100
The gist: Large language models are widely deployed as general-purpose problem solvers, making accurate confidence estimation critical for reliable use, and this work introduces capability calibration to
Key concepts
- Response Calibration
- This method assesses the certainty of a single response generated by the model. It checks if one specific answer is correct or incorrect. This is limited because LLMs are stochastic; it doesn't reflect how good the model is overall.
- Capability Calibration
- This targets the expected accuracy of the entire output distribution. It calculates the probability that any response sampled from the model's full set of possibilities will be correct for a given query. This provides a more robust measure of true model capability.
- Jensen Gap
- This quantifies the difference between response calibration and capability calibration. It is mathematically equivalent to the variance of response correctness under the model's output distribution, showing precisely how much information is lost when only using response-level confidence.
Terminology
Summary
Large language models are widely deployed as general-purpose problem solvers, making accurate confidence estimation critical for reliable use, and this work introduces capability calibration to address its limitations. The core finding is that response-level confidence fails to reflect underlying model capability due to the stochastic nature of modern LLM decoding, leading to the introduction of a framework that targets the model’s expected accuracy on a query.
The Gist
Capability calibration targets the model’s expected accuracy on a query, which is formally defined as the probability that a response sampled from the model’s output distribution conditioned on x is correct.
Defining Calibration Targets and Their Divergence
The paper distinguishes between two primary calibration types: response calibration and capability calibration. Response calibration assesses the correctness of a single generated output,
while capability calibration calibrates confidence against the expected accuracy of fθ's output distribution.
The authors formally show that these two notions differ both theoretically and empirically. Specifically, the optimal confidence scores are distinct: the optimal response-calibrated confidence is a binary indicator of a response’s correctness,
whereas the optimal capability-calibrated confidence is a continuous probability representing capability.
This divergence is quantified by the Jensen Gap, which is equivalent to the variance of response correctness under the model’s output distribution.
The Capability Calibration Framework
Capability calibration targets the model's expected accuracy, defined as:
- The target expected accuracy µ(x, fθ) is defined as:
µ(x, fθ) ≜ Pyˆ∼fθ(·x)[C(x, yˆ) = 1] = Eyˆ∼fθ(·x)[C(x, yˆ)].
- The perfect calibration for capability calibration is s∗ = µ(x, fθ).
To evaluate this target empirically when it is not directly observable, the authors develop an empirical framework that approximates it through repeated sampling. They estimate the expected accuracy by sampling keval responses
and computing the empirical mean:
µˆ = c/keval.
The choice of sample size guidance suggests using keval = 100 to balance cost and reliability, as this provides a stable per-query estimate of µ while keeping evaluation cost tractable.
Confidence Estimation Methods
The paper investigates several methods for producing calibrated scores, categorizing them into training-free and training-based approaches. The methods evaluated include:
-
Uniform random baseline: Sampling a confidence score from U(0, 1).
-
Response consistency: Estimating confidence by measuring
the fraction that agree with the majority prediction
across multiple sampled responses, though this incurs a higher computational cost as it requireskc forward passes per query.
-
Verbalized confidence: Instructing the LLM to report a probability in [0, 1] regarding its ability to answer the query correctly, focusing on
confidence in the query itself
rather than just the response. -
Probing LLMs’ hidden states: Training linear probes on activations (e.g., mean-pooled activations) to predict query-level confidence, which is noted as a
white-box method requiring open-weight models.
Practical Applications and Utility
Capability calibration enables practical applications by leveraging the model’s expected accuracy for downstream tasks. The authors demonstrate its utility in two representative settings:
-
Pass@k prediction: Capability-calibrated confidence scores are used to predict the pass@k success rate of individual queries
without extensive sampling,
simulating performance using Oracle Capability-Calibrated (Oracle-CC) confidence, whichsimulates groundtruth performance almost perfectly.
-
Inference budget allocation: Confidence estimates guide the allocation of computational resources across queries, where
higher confidence requiring fewer resources
is utilized. The greedy algorithm for this allocation relies directly on the expected accuracy formulated in Equation (4), allowing for an analytical estimation of the marginal improvement (gain) in success rate achieved by adding a single sample.
Empirical Results and Trade-offs
Experiments across three LLMs (Olmo-3-7B-Instruct, Qwen3-8B, gpt-oss20b) and seven datasets (including TriviaQA, GSM8K, MATH, AIME25) show that probing on model activations offers a favorable tradeoff between computational cost and confidence estimation performance.
Specifically, the best method is often Probe,
which performs well under in-domain settings. The results confirm that capability calibration leads to improved performance over baselines in both pass@k simulation and inference budget allocation, with Oracle CC achieving near-perfect simulation accuracy. However, the paper notes that while probing is effective for certain domains, developing generalizable methods for capability calibration remains an important direction.
Conclusion
The work formalizes capability calibration as a necessary refinement over response calibration due to LLM stochasticity.
Improvements for AI systems
As a fastidious and diligent researcher, I have analyzed this paper, On Calibration of Large Language Models: From Response To Capability,
which formally introduces and validates capability calibration as a superior confidence estimation target over traditional response-level calibration in stochastic LLMs.
Based on the findings presented in the paper, here are the specific improvements we can implement to AI systems and what those improved systems can achieve:
)
)
-
Implement a new confidence estimation layer called
Capability Calibration Estimator
that replaces or augments existing response-level confidence scores (e.g., token probabilities or response consistency metrics). -
This estimator must be trained using methods identified as effective, specifically recommending the use of training linear probes on LLM hidden states (mean-pooled activations) as a practical, low-cost white-box method for generating query-level confidence scores.
-
The system will use this Capability Calibration Estimator to generate a continuous probability score, denoted as the
expected accuracy
or capability calibration score, rather than a binary correctness indicator for any single response.
)
)
- Instead of relying on the simple Brier Score to evaluate calibration quality, the system must utilize the derived loss function that decouples output variance from response calibration:
L capability Brier = E[L response Brier] - Variance(C(x, yˆ)).
- This allows for a more robust assessment of model capability by explicitly penalizing stochasticity in the model's output distribution, providing a metric that is theoretically sounder than simple response-level metrics.
)
-
For resource allocation tasks (e.g.,
best-of-k
inference), integrate the proposed capability calibration confidence score directly into the greedy budget allocation algorithm (Damani et al., 2024). -
The improved system will use the formula derived in Equation (20): Gain i = p i(1 - p i) / k, where the probability of success is defined by capability calibration confidence. This ensures computational resources are allocated to queries where the model has a genuinely high expected accuracy, leading to significantly better performance than uniform allocation.
)
-
For real-time decision-making and deployment scenarios, implement an
Inference Budget Allocator
that uses the capability calibration score (e.g., derived from Probe-MATH or Verbalized Confidence) as a threshold for system behavior. -
The improved system will perform:
a) Selective Prediction: If the confidence score falls below a critical threshold, the system automatically abstains from providing an answer or triggers a request for human oversight (as suggested by Kamath et al., 2020).
b) Active Query Refinement: If confidence is low, the system performs query rewriting to ask for clarification before attempting a final response.
)
-
For large-scale knowledge retrieval and evaluation pipelines, deploy a
Pass@k Simulator
module that utilizes the Oracle Capability-Calibrated Confidence score (Equation 4). -
This module will allow researchers to simulate the performance curve of an LLM over a dataset without requiring multiple expensive sampling passes or assuming prior distributions, providing an unbiased estimation of how well the model can solve a query at various recall levels (k).
)
-
In
Label-free Benchmarking
andCurriculum Learning
pipelines, use the capability calibration score as a proxy for instance difficulty. -
This allows for automated sorting of training data to prioritize difficult instances or to evaluate model generalization on unseen domains by comparing the calibrated confidence estimates across different task types (Factual Knowledge vs. Mathematical Reasoning), leading to more efficient and targeted model training strategies.
Sources
- gpt-oss-120b & gpt-oss-20b Model Card
- A Comprehensive Survey on the Trustworthiness of Large Language Models in Healthcare
- Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
- Monte Carlo Temperature: a robust sampling strategy for LLM's uncertainty quantification methods
- No Answer Needed: Predicting LLM Answer Accuracy from Question-Only Linear Probes
- FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance
- Query-Level Uncertainty in Large Language Models
- Evaluating Large Language Models Trained on Code
- Training Verifiers to Solve Math Word Problems
- Learning How Hard to Think: Input-Adaptive Allocation of LM Computation
- Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty
- Uncertainty Estimates of Predictions via a General Bias-Variance Decomposition
- SimpleQA Verified: A Reliable Factuality Benchmark to Measure Parametric Knowledge
- A Baseline for Detecting Misclassified and Out-of-Distribution Examples in Neural Networks
- Measuring Massive Multitask Language Understanding
- The Curious Case of Neural Text Degeneration
- A Survey of Uncertainty Estimation in LLMs: Theory Meets Practice
- TrustLLM: Trustworthiness in Large Language Models
- LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative Fusion
- TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering