Stacked conformal prediction

summary

Video file (mp4)

The gist

Pr(Yn+1 ∈ C(α)n+1) ≥ 1 − α, where C(α)n+1 = y ∈ R: R(Yn+1)n+1 ≤ R(⌈(1−α)(n+1)⌉)(Yn+1).

In short

This episode discusses 'Stacked conformal prediction,' a method for creating reliable and computationally efficient prediction intervals. The hosts explain how stacking multiple base models and applying conformal prediction can yield tighter, valid intervals without needing to waste data on separate calibration sets.

Key concepts

Conformal Prediction
A statistical technique used to create reliable prediction intervals (ranges where an answer is expected). It ensures that the predicted range has a guaranteed coverage level, providing a measure of uncertainty for model outputs.
Stacked Models
An ensemble method where multiple base learners are trained first. A meta-learner is then placed on top to combine the predictions from these base models, potentially increasing predictive power.
Sherman-Morrison Update
A classical matrix formula used to efficiently update the inverse of a Gram matrix. This allows for computationally manageable updates when adding test points, avoiding expensive full retraining.
Full Conformal Prediction
The ideal form of conformal prediction where every data point is used for both training the model and calibrating the prediction intervals, eliminating data waste.

Terminology used across episodes

This episode discusses

The paper

Stacked conformal prediction · Read on arXiv

Insper Institute of Education and Research

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Stacked conformal prediction".

Jane: The paper was written by Paulo C. Marques F from Insper Institute of Education and Research.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title and Core Idea: Tom: Welcome back to the show, everyone. Today we're digging into a fresh arXiv paper called "Stacked conformal prediction," and I've got Jane here with me. Jane, before we bring in the rest of the crew, what's the one-line pitch?

Jane: Tom, it's about making prediction intervals—you know, the range where a model says the answer will fall—both more reliable and cheaper to compute. The authors, Paulo C. Marques F. from Insper, found a way to combine multiple models into a stack and then wrap that whole thing in conformal prediction without needing to hold out a separate chunk of data just for calibration.

Tom: And that separate calibration set thing is the part that always bugged me. You split your data, you train on less, you calibrate on less, and you just hope you didn't waste anything.

Jane: Exactly. Standard inductive conformal prediction forces that trade-off. This paper says, what if we stack models—train several base learners, then a meta-learner on top—and conformalize the whole stack in a way that uses all the data for both training and calibration?

Tom: But hold on, doesn't full conformal prediction usually mean retraining the model a million times for every new test point? That's the computational nightmare that made people invent the split version in the first place.

Jane: That's the clever bit. The meta-learner at the top of the stack is simple—they specialize it to multiple linear regression—and that simplicity lets you do the retraining efficiently using a classical matrix formula called the Sherman-Morrison update. So you get the statistical benefits of full conformal prediction without the computational explosion.

Tom: So the title "Stacked conformal prediction" is literally about stacking models and then conformalizing the stack. And the author is betting that the meta-learner being simple is the key that unlocks everything.

Jane: Right. And the paper proves that if you build the stack symmetrically—which is an oracle version you can't actually run—you get exact marginal validity. Then they show that if the base models are stable, the feasible version gets approximate validity. It's a nice theoretical bridge.

Tom: I love that they're honest about the oracle construct. They build this perfect symmetric stack in theory, prove the exchangeability property, and then say, okay, now let's make it real and see how much we lose.

Jane: And the empirical results suggest you don't lose much. We'll get into the numbers in a bit, but I'm excited to hear what Lu and Meng think about the practical side.

Tom: Before we bring them in, let me just say—this feels like one of those papers that could make conformal prediction actually usable in production settings where you can't afford to waste data.

Jane: Exactly. And that's the hook for our next segment, where we'll walk through the actual method and the theory behind it.

Methodology and Theory: Tom: So we're back with "Stacked conformal prediction," and I want to bring in Lu from Tsinghua. Lu, Jane and I were just talking about the oracle stack and the feasible stack. What's the real meat of the theory here?

Lu: Thanks, Tom. The key idea is that they construct the stack so that the second-level data—the predictions from the base learners paired with the true responses—become exchangeable. That's the property conformal prediction needs. The oracle version includes the future point in the training folds, which makes everything symmetric, and they prove that exchangeability transfers up the stack.

Jane: And that's Proposition one right? The second-level pairs are exchangeable if the base learners treat their training data symmetrically.

Lu: Precisely. Then Proposition two uses a standard conformal argument to get exact marginal validity for the prediction set. The coverage guarantee is at least one minus alpha, as long as you pick the right quantile of the conformity scores.

Tom: But the oracle version is impossible in practice because you don't know the future response when you're training. So they break the symmetry by pulling the future point out. What's the damage?

Lu: That's Proposition three. They show that if the base learners are stable—meaning the conformity scores don't change too much when you remove one point—then the coverage guarantee degrades gracefully. You lose a bit of coverage, controlled by a stability probability and a small epsilon term.

Meng: Lu, I'm an engineer, so let me ask the practical question. How do you actually compute the conformity scores without retraining everything from scratch?

Lu: Great question, Meng. That's where the Sherman-Morrison formula comes in. For multiple linear regression, adding a hypothetical test point to the training set changes the design matrix by a rank-one update. The Sherman-Morrison formula lets you update the inverse of the Gram matrix in O(M squared) time instead of O(M cubed) for a full inversion.

Meng: So for each candidate value of the future response, you're doing a cheap update, computing residuals, and checking whether the score falls below the quantile. That's what Algorithms one and two in the paper do.

Lu: Exactly. And the residuals are handled cleverly—they use leave-one-out style residuals to make the prediction intervals adaptive to heteroscedasticity, which is when the noise varies across the feature space.

Jane: I want to emphasize that this is full conformal prediction, not the split version. So every training point is used for both fitting and calibration. No data wasted.

Tom: And the cost is manageable because the meta-learner is just linear regression. That's the whole trick—stacking gives you the predictive power of complex base learners, and the simple top layer gives you tractable conformalization.

Lu: Right. And the paper is honest that this is approximate validity for the feasible stack, not exact. But the approximation is controlled by a stability assumption, which is reasonable for many real-world models.

Meng: I'm curious about the actual numbers. How does it perform on real data?

Tom: That's exactly what we're going to dig into next. We've got the California housing and Ames housing results to look at.

Experiments and Results: Tom: We're back with "Stacked conformal prediction," and now it's time to talk about whether this thing actually works. Meng, you asked about the numbers—let's get into them.

Meng: Yeah, I want to see if the approximate validity holds up in practice and whether the intervals are actually tighter than the standard alternative.

Jane: So they tested on two datasets. California housing has about twenty thousand census tracts with eight predictors, and the response is median house value. Ames housing has about three thousand houses with eighty predictors, and the response is sale price.

Lu: And they compared against conformalized quantile regression, or CQR, which is a popular inductive method. For CQR they used Quantile Random Forests as the underlying model.

Tom: The headline result, at least for me, is that the stacked conformal prediction intervals are consistently narrower. Look at the California dataset at the ninety percent nominal coverage level—the median interval width is about one hundred nineteen thousand dollars for the stacked method, but CQR gives you one hundred fifty-six thousand six hundred.

Meng: That's a thirty percent reduction in median width. That's substantial.

Jane: And the empirical coverage is right where it should be. For California at ninety percent nominal, they got eighty-nine point nine percent empirical coverage. For Ames at ninety percent, they got ninety-one point one percent. So you're not sacrificing validity to get those narrower intervals.

Lu: The pattern holds across all the coverage levels they tested—eighty percent, eighty-five percent, ninety percent. The stacked method always has comparable or better coverage and consistently shorter intervals.

Meng: What about the computational cost? The paper claims the Sherman-Morrison updates make it manageable, but what does that mean in practice?

Jane: The paper doesn't give wall-clock timings, but the algorithmic complexity is clear. For each test point, you're doing a binary search over the response range, and each step involves a rank-one update and residual computation. The base learners—Random Forests and CatBoost—are trained once on the full training set.

Tom: And that's the key advantage over full conformal prediction with a complex model. You'd have to retrain the Random Forest for every candidate response value, which is insane. Here, the expensive models are trained once, and the cheap linear meta-learner does the heavy lifting during conformalization.

Lu: I also want to point out that the residual construction in Algorithm two is what makes the intervals adaptive. They're not just using absolute residuals—they're using leave-one-out style residuals that account for the influence of each training point. That's why the intervals can be wider in noisy regions and narrower in clean regions.

Meng: So the practical takeaway is that you get tighter intervals, valid coverage, and a computational path that's actually feasible. That's a strong combination.

Jane: And it's worth noting that the paper includes code in R, Python, and C++ for full reproduction. That's a nice touch for people who want to try it themselves.

Tom: Before we wrap up, I want to get Lalam's take on where this could go next. Lalam, what's the big picture here?

Conclusion: Tom: So we've been talking about "Stacked conformal prediction" all episode, and I think we've covered the theory, the method, and the results. Let's bring in Lalam to help us see the forest through the trees.

Lalam: Thanks, Tom. What excites me about this paper is that it removes a practical barrier to using conformal prediction in real systems. The split version wastes data, and the full version is computationally prohibitive for complex models. This stacked approach threads the needle.

Jane: And the fact that the intervals are narrower means you're getting more precise answers, not just valid ones. That's what users actually care about—they want a tight range they can act on.

Lu: The theoretical contribution is also solid. The exchangeability transfer through the stack is elegant, and the stability-based approximation is a clean way to handle the feasible case.

Meng: From an engineering standpoint, the Sherman-Morrison trick is the kind of thing that makes a method deployable. You can integrate this into a prediction service without worrying about retraining costs exploding.

Tom: And the authors mention that the framework extends beyond regression—classification is a natural next step. The k-nearest neighbors meta-learner could work well there.

Lalam: I'd add that this could have real cultural impact in domains where decisions are made under uncertainty—healthcare, finance, climate modeling. Tighter, valid prediction intervals mean better risk assessment and more informed decisions.

Jane: It's also democratizing in a way. The code is open source, the method is model-agnostic at the base level, and the computational requirements are modest. Smaller teams can adopt this without needing massive compute budgets.

Tom: Alright, let's wrap this up. "Stacked conformal prediction" by Paulo C. Marques F. gives us a way to conformalize stacked ensembles with approximate validity, tighter intervals than the standard inductive approach, and manageable computational cost.

Meng: And the empirical results back it up—valid coverage and narrower intervals on both California housing and Ames housing.

Lu: It's a nice piece of work that bridges theory and practice. I'm looking forward to seeing extensions and follow-ups.

Lalam: And I'm looking forward to seeing it applied in real systems where uncertainty quantification matters.

Jane: Great discussion, everyone. We'll be back next time with another paper, but for now, this is Tom and Jane signing off on "Stacked conformal prediction."

Tom: Thanks for listening, and keep an eye on the arXiv. There's always something new to learn.

More episodes

← Home