V-ECE: Estimating General Expected Calibration Errors

arXiv:2602.24230 · stat.ML, cs.LG · Submitted 2026-02-27 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "V-ECE: Estimating General Expected Calibration Errors".

Tom: The gist: This paper introduces an extension to variational frameworks for estimating calibration errors that can cover any binary or multiclass Lp calibration error,

Jane: First, who's behind it and why it matters.

Title and authors: Tom: So we're talking about the paper called V-ECE: Estimating General Expected Calibration Errors. The title tells you right away that this work is aiming to tackle calibration error in a much broader way than what we usually see.

Jane: Exactly, Tom. It’s not just looking at one specific type of error anymore, but trying to cover any Lp calibration error, no matter how complex the loss function might be. It’s about making the tools more general for checking if our models are actually reliable.

Lu: What's interesting is that this paper extends an existing variational framework. They’re taking an idea from Berta et al. and making it work for any binary or multiclass Lp calibration error, which is a big step toward generalization across different model setups.

Meng: So, when you look at the general scope, what does that mean practically for someone building a system? Are we talking about something that just gives us one number to check our model's health?

Tom: Well, it’s more than just one number. The paper shows how this method can separate over-confidence from under-confidence, which is a real distinction in how we interpret model behavior. This is important because sometimes the error isn't just about being slightly too sure or slightly too unsure.

Jane: It’s about getting a clearer picture of where the model is actually failing, rather than just seeing one big error score that hides both issues together. That separation is key for debugging these systems.

Lu: And they achieve this by using cross-validation to estimate the recalibration function, which helps them avoid overestimating the true calibration error in expectation. This technique bypasses some of the pitfalls we see in other non-variational approaches that tend to be biased.

Meng: Avoiding overestimation is a big deal for me because I’ve seen how those methods can sometimes inflate the error score just because they are too optimistic about their own estimate. If this method gives us a lower bound, that’s much more trustworthy.

The paper's summary: Tom: So, what is the main mechanism here? Basically, they’re using a variational approach to break down the total risk of the classifier into two parts: one related to how well the model predicts things versus how well it matches what actually happens.

Jane: That decomposition relies on proper losses first, which we know are easier to handle, but this paper pushes that idea further by showing how you can adapt it for any Lp calibration error, not just the ones that fit a standard definition.

Lu: They achieve this by allowing the entropy function and thus the proper loss to change depending on what input data f(X) we are looking at. This is what lets them recover those divergences that aren't caused by a single, fixed proper loss function across all inputs.

Meng: That sounds mathematically intensive. From an engineering standpoint, how does changing the entropy function for every single prediction actually translate into something useful for deployment? Does it complicate the runtime too much?

Tom: It complicates things initially, but the paper shows that when you use this approach to estimate the Lp calibration error, like CE p(f), you can define a specific proper loss function that captures exactly what we’re trying to measure.

Jane: They show that for any p greater than or equal to one, they can recover the Lp calibration error by looking at the difference between the risk of your initial model and the risk of a recalibrated version, defined as CE p(f) = one over n sum i=one n f(X i)(f(X i), Y i), where you average this over several cross-validation estimates <ref:2602.24230#pg1>.

Lu: And they confirm that this process leads to the recovery of the Lp calibration error as CE p(f) = E

f(X)(f(X), Y) - f(X)(f(X), Y): (three) <ref:2602.24230#pg1>. This is the core mathematical result for extending proper errors to Lp errors.

Meng: So, it’s essentially a complex way to measure how far our current model's prediction risk is from the best possible risk achievable by some recalibration function. That sounds like a lot of computation for real-time use.

The paper's improvements: Tom: The paper points out two major advantages over existing methods. First, using cross-validation to estimate that recalibration function guarantees that the calibration error is not overestimated in expectation.

Jane: That’s a big deal because, as we discussed, many older methods could just give you an inflated error score simply by being too optimistic about their estimation of the recalibration function. This cross-validation approach provides that lower bound we talked about earlier.

Lu: They also demonstrate empirically that this method converges to the true calibration error faster than classical binning-based estimators, which is a strong practical finding when you’re trying to assess model performance quickly.

Meng: Faster convergence is good for development cycles. If we can get a reliable estimate of calibration error quicker, we can iterate on our models more rapidly instead of waiting for slow, traditional checks.

Tom: And they also introduce ways to isolate the over- and under-confidence aspects by defining specific proper losses like f(X),+(p, y) and f(X),-(p, y) for the binary case.

Jane: That refinement is really insightful. They found that these specific losses help them detect no under-confidence in the over-confident scenario, which suggests that calibration error is completely determined by that over-confident part of the model.

Lu: It means they’ve refined the analysis to be more nuanced; instead of just a single number, you can dissect *why* the calibration is off based on these confidence measurements.

Meng: So, it moves from just saying "the error is high" to saying "the error is high because we are over-confident in this specific way," which gives us much better direction for fixing the model.

Conclusion: Tom: We’ve covered a lot about V-ECE: Estimating General Expected Calibration Errors. It really shows how to extend variational frameworks to handle any Lp calibration error and provides a robust way to get lower bounds on true error using cross-validation.

Jane: So, the main implication for us is that we can stop relying solely on binning methods like ECE when we need more precision, because this framework offers a path to estimate Lp errors consistently across different loss functions.

Lu: The future work seems to be applying this framework even more deeply, exploring how it might interact with other model architectures or complex data structures. It’s open-ended possibilities for what we can explore next.

Meng: For me, I think the practical path forward is integrating these refined loss definitions into our standard validation pipelines so that developers get that faster convergence we talked about earlier when testing new models.

Lalam: I see the potential here to make our own internal evaluation tools much smarter, allowing us to measure model calibration with this level of detail, which could really improve how we assess the quality across the board.

INRIA 2Ecole Normale Supérieure, PSL Research University 3UC Berkeley

stat.ML, cs.LG

Submitted: 2026-02-27

Updated: 2026-10-08

Comments: Re-worked version with a new metric benchmark and new mathematical results on the bias of the estimator

Code: https://github.com/dholzmueller/probmetrics

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 89/100

The gist: The gist: This paper introduces an extension to variational frameworks for estimating calibration errors that can cover any binary or multiclass Lp calibration error, offering benefits like

Key concepts

Proper Calibration Error
This is the fundamental risk decomposition showing how total risk relates to an initial model's risk minus the minimum achievable risk after applying an optimal recalibration function. It helps isolate where calibration issues originate.
Lp Calibration Error (CE||·||p)
This generalizes calibration error beyond standard metrics by allowing the loss function and entropy to change for every output prediction. It is estimated by fitting a recalibration function using cross-validation and averaging the results across different data splits.
Over- and Under-Confidence Analysis
By adjusting the proper loss function, researchers can isolate components of calibration error related to over-confidence (being too certain) or under-confidence (being too uncertain). This reveals that in over-confident scenarios, the total error is driven by the confident component.

Terminology

Summary

The gist: This paper introduces an extension to variational frameworks for estimating calibration errors that can cover any binary or multiclass Lp calibration error, offering benefits like separating over- and under-confidence and avoiding overestimation through cross-validation; <ref:2602.24230#pg2>

Estimating Proper Calibration Errors

Proper calibration errors arise from the decomposition of the risk of a classifier when using a proper loss function, specifically as shown by Bröcker (2009) where E[l(f(X), Y)] = E[dl(f(X), C)] + E[el(C)] <ref:2602.24230#pg4>. A variational estimator for this proper calibration error, as described by Berta et al. (2025a), decomposes the error into CEdl(f) = E[l(f(X), Y)]−min g∈H E[l(g◦f(X), Y)] <ref:2602.24230#pg5>. This allows for the estimation of proper calibration error by evaluating the risk of the initial model minus the risk of an estimated recalibration function ˆg ◦ f, defined as CEc dl(f) = 1/n Σl(f(Xi), Yi) −l(ˆg ◦ f(Xi), Yi) <ref:2602.24230#pg7>.

Estimating Lp Calibration Errors

The paper extends the proper calibration error estimator to any Lp calibration error, CE∥·∥p (f), for any p ≥ 1, by allowing the entropy function H (and thus the proper loss) to change for every f(X) <ref:2602.24230#pg8>. Proposition 1 demonstrates that this leads to the recovery of Lp calibration error as CE∥·∥p (f) = E[lf(X)(f(X), Y) − lf(X)(g⋆◦ f(X), Y)] <ref:2602.24230#pg9>. The proper loss induced is given by lf(X)(z, Y) = ⟨∇z∥z − f(X)∥p, z − Y⟩ - z − f(X)∥p for every z ≠ f(X) <ref:2602.24230#pg10>. In practice, the Lp calibration error of a classifier f is estimated by fitting ˆg with k-fold cross-validation and computing CEc∥·∥p (f)j = -1/Ij Σlf(Xi)(ˆg ◦ f(Xi), Yi) and averaging the k estimates.

Lower Bound with Cross-Validation

A key benefit of the proposed procedure is obtaining a lower bound on the true calibration error using cross-validation, which circumvents the pitfall of overfitting the recalibration function ˆg. The benefit of using cross-validation is that learning ˆg and estimating the calibration error with different samples guarantees that the calibration error is not over-estimated (in expectation). For every function ˆgi learned on cross-validation split i, E[l(ˆgi ◦ f(X), Y)] ≥ E[l(g⋆ ◦ f(X), Y)] because g⋆ minimizes E[l(g ◦ f(X), Y)].

Estimating Over- and Under-Confidence

The framework can be refined to isolate over- and under-confidence by adjusting the proper loss, specifically defining lf(X),+(p, y) and lf(X),−(p, y) for the binary case. In the multiclass setting, a one-versus-rest fashion is adopted to evaluate over- or under-confidence of the top class prediction using these losses. Experiments show that using these losses provides a refined analysis, detecting no under-confidence in the over-confident scenario and revealing that the calibration error is fully determined by the over-confident component.

Experimental Results

Experiments compare various classifiers, including TabICLv2 and RealTabPFN-2.5, against classical estimators like ECE (Binning) and Isotonic regression. Using cross-validation returns a lower bound on the true CE·. The state-of-the-art classifiers TabICLv2 and RealTabPFN-2.5 recover the most calibration error in binary and multiclass settings.

Classifier Selection

The paper evaluates several models, including Tabular foundation models like RealTabPFN-2.5 (Grinsztajn et al., 2025) and TabICLv2 (Qu et al., 2026), as well as gradient-boosted decision trees such as CatBoost and LightGBM. The recommendation is to use the logitinitialized CatBoost classifier as the default model in their package.

Hardware Considerations

The experiments ran tabular foundation models on GPUs (NVIDIA V100) and other models on CPUs (Cascade Lake Intel Xeon 5217 with 8 cores). The trade-off noted is that while using cross-validation allows for post-hoc calibration, it incurs a higher runtime due to cross-validation.

Algorithmic Procedure

The proposed algorithm is summarized in Algorithm 1, which involves partitioning data into k folds and training a classifier g(j) on the training set of each fold. The final calibration error is aggregated across folds as CE[dl = X/k Σj=1 Ij n CE[dl(j)].

General Distance Estimation

Proposition 2 generalizes the theorem to show that general distances d(p, q) on the simplex can be written in a variational form whenever all functions Dp(q) = d(p, q) are convex and minimized at p. This leads to a proper loss satisfying EX[Df(X)(C)] = E[lf(X)(f(X), Y)] − inf g∈H E[lf(X)(g◦f(X), Y)].

Multiclass Extension Details

In the multiclass case, the L1 calibration error is simply the sum of the binary L1 calibration error of each class against the rest, sometimes called “pairwise calibration error” E[∥f(X) − C∥1] = Pk i=1 E[f(X)i − Ci] <ref:2602.24230#pg10>. The results for CE∥·∥2 are shown in Figure 3. The true calibration error is therefore CE∥·∥p = E[∥U − g⋆(U)∥p].

Over-confident Under-confident Analysis

Figure 2 illustrates different simulated mis-calibration scenarios, including over-confident predictions, under-confident predictions, or a mix of both. Table 2 shows the CE recovered for the different over- and under-confident setups in Figure 2. These losses provide a refined analysis, detecting no under-confidence in the over-confident scenario and revealing that the calibration error is fully determined by the over-confident component.

Synthetic Experiments

In synthetic experiments, for binary classification, u ∈ [0, 1], three settings are tested: perfectly calibrated case g⋆(u) = u, over-confident case g⋆(u) = sigmoid(0.4 · logit(u) + 0.3), and the shifted case g⋆(u) = min(1, u + ε) with ε = 0.02. In the multiclass case, u ∈ ∆d, experiments test perfectly calibrated case g⋆(u) = u, overconfident case g⋆(u) = sigmoid(0.3 · log(u)), and under-confident case g⋆(u) = sigmoid(2 · log(u)).

Improvements for AI systems

  1. Improved calibration estimation via variational framework: The method can separate over- and under-confidence by using a variational estimator that avoids overestimation, which provides a more accurate measure of true calibration error than classical binning methods like ECE, which are noted as being biased and inconsistent.

  2. Accurate Lp calibration error estimation: The system can estimate any Lp calibration error, including those induced by non-proper losses, using the derived proper loss formulation: For p ≥ 1, z ∈ ∆k, Y ∈ Y, define lf(X)(z, Y):= 1z̸=f(X)⟨∇z∥z − f(X)∥p−1∥z−f(X)∥p−1p, with element-wise power and sign in the numerator.

  3. Guaranteed lower bound estimation: By using cross-validation to estimate the recalibration function without overestimating, the system can ensure that it yields a lower bound on the true calibration error (in expectation) by evaluating risk on hold-out splits.

  4. Refined over- and under-confidence detection: The framework allows for isolating these effects by defining specific proper losses, such as "l f(X),+(p, y):= l(1f(X)>1/2 min[p, f(X)] + 1f(X)<1/2 max[p, f(X)] + 1f(X)=1/2/2, y) and l f(X),−."

  5. Faster convergence to true error: The approach is empirically shown to converge faster than classical, binning-based, estimators, suggesting quicker calibration assessment in practice compared to traditional methods.

Abstract

In probabilistic classification, calibration error (CE) measures the average divergence of predicted probabilities f(X) from P(Yf(X)), the true class distribution for that predicted probability. While being a useful diagnostic tool, it is hard to estimate: popular binning-based estimators are often inconsistent and scale poorly beyond two classes. Recent work rewrites the CE as the excess risk of a model compared to the best recalibration of its own predictions, measured with a proper loss. However, this only works for Bregman-divergence-based calibration errors like the squared error, excluding the more popular L 1-distance-based CE. We show that using prediction-dependent proper scores can alleviate this restriction, allowing us to estimate CEs with general convex divergences, including L p distances with closed-form losses in the binary and multiclass settings. To estimate the excess risk, we introduce a more accurate recalibrator that fits a residual to temperature scaling with gradient boosting. The resulting variational estimator, V-ECE, needs no bins or clusters and lower-bounds the true calibration error in expectation. On a benchmark of semi-synthetic tasks built from real classifiers, with known true CE, V-ECE is among the most accurate binary estimators for every calibration error and significantly outperforms all multiclass estimators. Our results are accompanied by additional theory on L p CE, estimator bias, and over- or under-confidence estimation.

Sources

Related papers