Marginally Useful: Formalizing the Information Gap in Conformal Prediction
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Marginally Useful: Formalizing the Information Gap in Conformal Prediction".
Jane: The paper was written by Peter Cotton from Microprediction.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back to the arXiv radio hour, everyone. Today we're digging into a paper with a title that's just deliciously pointed: "Marginally Useful: Formalizing the Information Gap in Conformal Prediction." Jane, I have to say, that title alone got me hooked.
Jane: Oh, absolutely, Tom. It's a pun and a thesis statement all at once. The authors are telling us that conformal prediction, this hugely popular tool for uncertainty quantification, is only "marginally" useful — and they mean that literally, because it gives you marginal coverage, not the conditional coverage you actually want.
Tom: So for our listeners who haven't been living in the conformal prediction world, can you break down what that means? What is this tool that everyone's so excited about?
Jane: Sure. Conformal prediction is a way to take any prediction model and wrap it in a guarantee. You get a set of possible outcomes — like an interval around your forecast — and the guarantee says the true value will fall inside that interval, say, ninety percent of the time, no matter what the underlying data distribution looks like. That's a really powerful promise.
Tom: And the paper's saying that promise is real, but it's being oversold. The guarantee averages over all possible inputs, right?
Jane: Exactly. It's a guarantee about the average future case, not about the specific case in front of you. So if you're predicting something where the uncertainty varies a lot depending on the input — some cases are easy, some are hard — the conformal interval will over-cover the easy ones and under-cover the hard ones, and the ninety percent number is just the average of those two failures.
Tom: That's a pretty damning critique. But the paper isn't just saying "this is bad," right? It's actually formalizing exactly what's lost. There's a precise formula for the gap.
Jane: That's the real contribution. They show that for a certain class of conformal predictors, the loss in forecast quality compared to a perfect oracle is exactly the mutual information between the residual — that's the prediction error — and the input. If your model's errors depend on the input, which they almost always do in practice, then you're leaving information on the table that no amount of conformal calibration can recover.
Tom: And that's the "information gap" in the title. It's not a vague hand-wave about limitations; it's an exact identity. I love when a critique comes with math that sharp.
Jane: Me too, Tom. And the title "Marginally Useful" captures it perfectly — the guarantee you get is marginal, and the usefulness is marginal too, unless you're careful about what you're actually asking for.
Tom: So we've got the setup. Next we need to talk about what this means for people actually using conformal prediction in practice, because I suspect the implications are going to ruffle some feathers.
Summary: Tom: We're back with "Marginally Useful: Formalizing the Information Gap in Conformal Prediction." Jane, I want to dig into the core result now, because the paper doesn't just critique — it gives us a formula. Can you walk us through what they actually proved?
Jane: So the setup is this: you have a location predictor, say a neural network that outputs a single number, and you're using conformal prediction to add uncertainty around that number. The paper considers two common ways of doing that — one uses the signed residual, the other uses the absolute value of the residual. And in both cases, they show that the predictive distribution you get is just the marginal distribution of the residual, shifted by your location prediction.
Tom: Meaning it uses the same error distribution everywhere, regardless of the input.
Jane: Precisely. And the oracle — the perfect predictor — would use the conditional error distribution, which varies with the input. The gap between those two, measured in terms of log-score regret, is exactly the mutual information between the residual and the input. That's their Proposition three and it's a beautiful, clean result.
Tom: And mutual information, for our listeners, is just a measure of how much knowing the input tells you about the error. If it's zero, the errors are independent of the input, and conformal prediction is perfectly fine.
Jane: Right. But when is it zero? Almost never in real problems. Heteroscedasticity — where the noise level changes with the input — is the norm, not the exception. So the gap is almost always positive, and it's not something you can tune away.
Tom: And that's the kicker. The paper shows that no recalibration that ignores the input can reduce this gap. You can adjust the shape of your distribution, you can fix the coverage level, but that mutual information term is untouchable unless you actually condition on the input.
Jane: Exactly. And there's a nice geometric way to think about it. They introduce this coverage-score plane. Coverage is one axis, forecast quality is the other. Conformal prediction moves you horizontally — it fixes your coverage — but it never moves you vertically. It never improves your actual forecast score.
Tom: So it's like fixing the fence but not improving the cattle. The coverage certificate is real, but it's not evidence that your forecast is any good.
Jane: That's the analogy, and it's exactly the point. The paper even has a proposition showing that two forecasters with identical conformal sets can have arbitrarily different log-scores. The coverage guarantee is completely orthogonal to forecast quality.
Tom: So the summary is: conformal prediction gives you a real guarantee, but it's the wrong guarantee for most forecasting questions, and the paper has now quantified exactly how wrong it can be. What do we do about it? That's where I want to go next.
Improvements: Tom: We're still with "Marginally Useful: Formalizing the Information Gap in Conformal Prediction." So Jane, the paper's been pretty critical so far. Does it offer any constructive paths forward? What are the improvements it suggests?
Jane: It does, and I think that's what makes this paper fair rather than just negative. The key insight is that the fix isn't to abandon conformal prediction — it's to stop using it for things it wasn't designed for. The paper is very clear: if your objective is coverage — like, "I need a set that contains the answer ninety-five percent of the time" — then conformal prediction is the right tool, full stop.
Tom: So it's about matching the tool to the question. What about when you actually want a distribution, not just a set?
Jane: Then you need to condition on the input. The paper points to several existing methods that do this — conformalized quantile regression, Mondrian conformal prediction that bins the input space, methods that explicitly model heteroscedasticity. These all work by leaving the single-shape class that the impossibility result applies to.
Tom: And the paper's contribution is showing why those methods work, not just that they work. Because they're all doing the same thing: they're reducing that mutual information term by conditioning on X.
Jane: Exactly. And there's a nice discussion of what recalibration can and cannot do. If you're using absolute residuals, you're paying an extra penalty for throwing away the sign information — that's a skewness term. Recalibration can fix that. But it can never touch the mutual information term. That requires actual modeling of the conditional error distribution.
Tom: So the improvement the paper suggests is really a change in mindset. Stop bolting conformal prediction onto a point predictor and expecting it to give you a full distribution. If you want a distribution, build one. If you want a coverage guarantee, use conformal prediction for what it's good at.
Jane: And the paper has this nice litmus test for practitioners. Ask yourself: is your loss a function of whether the truth lands in a region, or of where it lands? If it's the former, conformal prediction is your tool. If it's the latter, you need a proper probabilistic model.
Tom: That's a genuinely useful decision rule. And it's backed by experiments in the paper showing that even a simple adaptive volatility model beats all the conformal methods on proper scoring rules, while conformal methods win on coverage.
Jane: Right. And the time-series case is particularly interesting, because exchangeability — the assumption that makes conformal prediction work — basically never holds for time series. The paper shows how the guarantees weaken as you move from split conformal to adaptive methods, and how even those don't give you pointwise conditional coverage.
Tom: So the improvements are really about honesty and matching tools to objectives. We should probably wrap this up and think about what this means for the broader field.
Conclusion: Tom: We're wrapping up our discussion of "Marginally Useful: Formalizing the Information Gap in Conformal Prediction." Jane, give us the final summary — what should our listeners take away from this paper?
Jane: I'd say the core message is that conformal prediction is a real, valuable tool, but it's been oversold as a general-purpose uncertainty quantification method. The paper gives us an exact formula for what's lost: the mutual information between the residual and the input. That's the gap that no amount of conformal calibration can close.
Tom: And the practical takeaway is to ask yourself what you actually need. A coverage guarantee for a set, or a full predictive distribution? Those are different products, graded by different instruments.
Jane: Exactly. If you need a set that contains the answer with high probability — for selective prediction, anomaly detection, compliance — conformal prediction is excellent. If you need to know how confident the model is for a specific input, you need conditional modeling, and conformal prediction won't give you that for free.
Tom: The paper also makes a nice point about the time-series case, where the exchangeability assumption that underpins conformal prediction usually fails, and the guarantees weaken to long-run averages or regret bounds.
Jane: Right. And the experiments back it up — a simple adaptive volatility model beats all the conformal methods on proper scores, while conformal methods win on coverage. They're solving different problems.
Tom: So this paper is really a call for intellectual honesty in the field. Use the right tool for the right job, and don't mistake a coverage certificate for evidence of forecast quality.
Jane: And the title says it all — "Marginally Useful." The guarantee is marginal, and the usefulness is marginal, unless you're careful about what you're asking. It's a paper that should be required reading for anyone who's been told that conformal prediction is the way to "add uncertainty" to their models.
Tom: Well said, Jane. That's a wrap on this one. Next up, we've got a paper on something completely different — I can't wait to dig into it. Thanks for listening, everyone.
Jane: See you on the next episode.
Peter Cotton
Microprediction
q-fin.ST, math.ST, q-fin.MF, stat.ML, stat.TH
Submitted: 2026-08-21
Updated: 2026-08-24
Comments: 14 pages, 7 figures. Interactive demonstrations and code at https://conformalprediction.net
Code: https://github.com/microprediction/timemachines
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 66/100
The gist: The paper "Marginally Useful: Formalizing the Information Gap in Conformal Prediction" by Peter Cotton addresses a fundamental category error in how conformal prediction is often applied and
Terminology
Summary
The paper Marginally Useful: Formalizing the Information Gap in Conformal Prediction
by Peter Cotton addresses a fundamental category error in how conformal prediction is often applied and interpreted. The author states: "conformal prediction is a method for certifying the coverage of a set, and it is being asked to estimate a distribution. These are different kinds of object, graded by different instruments, and no amount of tuning turns one into the other."
The paper's central contribution is a new result presented in Section 5, called the residual-information gap. For a single-shape residual predictive system, the log-score regret relative to the oracle is exactly the mutual information I(R; X) between the residual and the input.
The author proves (Proposition 3) that for a fixed location predictor µ̂ with residual R = Y − µ̂(X), in the large-calibration limit:
-
(A) Signed-residual CPS: E[log q⋆(YX)] − E[log πX(Y)] = I(R; X)
-
(B) Absolute-residual intervals: E[log q⋆(YX)] − E[log hsym(R)] = I(R; X) + KL(r̄ ∥ hsym)
Both are ≥ 0. The term I(R; X) vanishes iff R ⊥ X
; the extra KL term vanishes iff the marginal residual law is symmetric.
Crucially, no recalibration that ignores X can reduce I(R; X): it can at best replace the shape by r̄, removing the skewness term in (5) and reducing case (B) to case (A). Reducing I(R; X) itself requires conditioning on X.
The paper establishes several key results:
Proposition 1 (Coverage does not constrain the log-score): The split-conformal set is a measurable function of the nonconformity scores si and of s(x, ·); it is not a functional of the predictive density f unless f is explicitly used in the score.
Consequently, "two forecasters with the same location µ̂... but different predictive densities f have identical conformal sets, and identical marginal coverage at every level α, while their expected log-scores E[log f(YX)] may differ by an arbitrary amount." The proof constructs a degenerate case where the conformal set is fixed but the log-score diverges to −∞ as σ → 0 or σ → ∞.
Proposition 2 (Re-leveling, two cases): In the large-calibration limit under exchangeability: (A) The signed-residual CPS predictive distribution is Πx(y) = G(y − µ̂(x)), the base location forecast re-leveled by the marginal signed-residual law.
(B) Ordinary split conformal at level α returns the symmetric interval µ̂(x) ± q̂α, whose density is the symmetrized marginal residual law hsym(z) = ½(g(z) + g(−z)).
The paper reviews three known limits of the marginal guarantee:
-
Marginal is not conditional, and conditional is impossible for free: Lei and Wasserman (2014) show
any band with non-trivial finite-sample conditional validity has infinite expected length at almost every point of a continuous distribution.
Foygel Barber et al. (2021) sharpen this:a conditionally valid Cn has E[leb(Cn(x))] = ∞ at almost all non-atomic x,
and even relaxed subgroup coverageis impossible to attain beyond the trivial solution
of inflating the marginal level to 1 − αδ. -
Validity is trivially satisfiable:
A predictor that returns the whole outcome space with probability 1 − α and the empty set otherwise satisfies (1) exactly and conveys nothing.
The same coverage attaches to excellent and terrible predictors;marginal validity is therefore not evidence of a good method; it certifies the fence, not the cattle.
-
Without exchangeability, even the marginal guarantee fails: The guarantee rests on exchangeability of the augmented sample; under covariate shift, label shift, or temporal dependence it does not hold.
The paper introduces a coverage–score plane (Section 6) with coordinates (C, S) where C = P(Y ∈ Cα(X)) − (1 − α) and S = E[log f(YX)]. The geometric reading: Residual split conformal is a horizontal projection. It maps a report (Cα base, f) to (Cα conf, f): it changes the set so that C → 0 while leaving the density f (and hence S) untouched... It moves a point left, never up.
Conformal predictive systems do have a vertical coordinate, but that coordinate is governed by the single-shape restriction of Proposition 3, capped at the oracle minus I(R; X).
For the time-series case (Section 7), the paper distinguishes three guarantees: (i) split conformal under exchangeability gives finite-sample marginal coverage; (ii) adaptive conformal inference gives long-run, time-average coverage control under distribution shift
; (iii) online conformal under arbitrary shifts gives regret-style or local-interval performance rather than the original exchangeable guarantee.
An experiment on a nonstationary series shows fixed split conformal collapses far below target once drift sets in,
while adaptive methods (ACI, conformal PID) recover coverage. However, a simple probabilistic volatility model (an EWMA-variance Gaussian) matches or beats them on the proper score and returns a full distribution.
The author's skaters
forecaster has the best non-oracle interval score (6.57) and CRPS (0.86) here, because its predicted spread tracks the changing volatility.
Notably, a normalized conformal wrap re-levels its coverage to 1 − α at an essentially identical interval score (6.58 against 6.57),
confirming that the conformal step re-levels; it does not add sharpness.
Section 8 addresses when coverage is the objective, providing a litmus test: Is your loss a function of whether the truth lands in a region, or of where it lands?
If the former, conformal prediction is sound for: selective prediction and risk control, retrieval and shortlisting, anomaly detection via conformal p-values, and compliance/long-run frequency control.
Section 9 discusses scope and constructive escapes. The strong claims target post-hoc, marginal, split conformal prediction with a residual score.
Methods that condition on X (Mondrian/binned CPS, conformalized quantile regression) buy sharpness or conditional coverage by conditioning on X or optimizing a proper objective, which is exactly what Propositions 3 say is required.
End-to-end conformal training deliberately differentiates through the conformalizer to minimize expected set size subject to coverage,
illustrating that to obtain sharpness you must optimize for it, rather than read it off the coverage certificate.
The information in set size (Correia et al., 2024) shows the expected size of a calibrated conformal set lower-bounds the conditional entropy H(YX),
complementary to Proposition 3.
The paper concludes: "Conformal prediction is not weak uncertainty quantification; it is exact uncertainty quantification for a set-valued coverage objective, and mistaking a set-coverage certificate for evidence of distributional quality is the category error."
Improvements for AI systems
Based on the paper, here are the specific improvements I can make to AI systems, along with what the improved system can do:
What I will do:
Instrument the AI system with a real-time computation of I(R;X) — the mutual information between the residual R = Y - (X) and the input X. This is computed via a nonparametric estimator (e.g., k-nearest-neighbor mutual information) on a held-out calibration set, alongside the existing log-score and coverage metrics.
What the improved system can do:
-
Quantify exactly how much predictive distributional information is lost by using a single-shape (location-only) forecaster.
-
Flag when the forecaster's uncertainty estimates are systematically overconfident or underconfident in specific regions of the input space, even when marginal coverage looks fine.
-
Automatically trigger a switch to a conditional (heteroscedastic) model when I(R;X) exceeds a user-defined threshold (e.g., 0.1 nats), preventing silent degradation on hard inputs.
After these changes, the AI system will:
-
Know its own blind spots — it can compute exactly how much information is lost by ignoring input-dependent residual structure.
-
Adapt its uncertainty estimates — it switches to conditional shapes or studentized scores when the residual-information gap is large.
-
Monitor both coverage and sharpness — it never lets a good coverage number hide a bad predictive density.
-
Survive nonstationarity — it detects exchangeability failures and switches to appropriate adaptive methods with honest guarantees.
-
Report honestly — it communicates what is guaranteed (marginal, conditional, long-run) and what is not, preventing costly misinterpretations.
These improvements are directly actionable, grounded in the paper's formal results (Propositions 1–3), and address the exact failure modes the paper identifies: marginal-not-conditional, coverage-not-quality, and exchangeability-not-guaranteed.
Abstract
Conformal prediction gives finite-sample, distribution-free marginal coverage for a set. The guarantee is real, and it is often misread as evidence of forecast quality. We separate the two with one decomposition, which we call the residual-information gap: for a single-shape residual predictive system, the log-score regret relative to the oracle is exactly the mutual information I(R;X) between the residual and the input. Conformalization re-levels coverage but cannot touch this quantity, because it is a property of the predictor's shape class and not of calibration; no recalibration that ignores X reduces it. The familiar cautions about conformal prediction follow as context: marginal coverage is not conditional, validity is insensitive to sharpness, and the guarantee needs exchangeability.
Sources
- Conformal Prediction for Time Series with Modern Hopfield Networks
- Conformal Prediction and Trustworthy AI
- Conditional Coverage Diagnostics for Conformal Prediction
- An Information Theoretic Perspective on Conformal Prediction
- A Post-Processing Conformal Prediction Approach for Conditional Coverage via Pivotal Scores
- Pitfalls of Conformal Predictions for Medical Image Classification
- Questioning the Coverage-Length Metric in Conformal Prediction: When Shorter Intervals Are Not Better
- Learning Optimal Conformal Classifiers
- CRPS-Optimal Binning for Univariate Conformal Regression