Confidence intervals for the random forest generalization error
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Confidence intervals for the random forest generalization error".
Jane: The paper was written by Paulo C. Marques F from Insper Institute of Education and Research.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title and Authors: Tom: Welcome back to the show, everyone! Today we’re digging into a paper that’s got me genuinely excited — it’s called “Confidence intervals for the random forest generalization error,” and it’s by Paulo C. Marques F. from Insper Institute in Brazil. Jane, I have to say, the title alone sounds like it’s solving a problem I didn’t even know I had.
Jane: Oh, Tom, it’s such a good one. So, when you train a random forest — you know, that classic machine learning model that combines hundreds of decision trees — you get a number called the out-of-bag error, which is basically a free estimate of how well your model will do on new data. But it’s just a single number. There’s no sense of how confident you should be in it.
Tom: Right, and that’s the gap this paper fills. The author shows that the same byproducts you already get from training the forest — the out-of-bag information, the predictions from each tree — can be bootstrapped to give you a confidence interval around that error. No extra training, no splitting your data into train and test sets.
Jane: Exactly. And that’s huge for practitioners. I mean, when you’re building a model, you don’t just want to know “my error is five percent.” You want to know “my error is five percent, plus or minus something.” This gives you that something, almost for free.
Tom: And it’s not just for regression — it works for classification too. The paper runs simulations on both types of problems and shows the confidence intervals actually cover the true error at the rates you’d expect. Like, for a ninety-five percent interval, you get coverage around ninety-four percent or ninety-three percent. That’s solid.
Jane: The author also tests it on real datasets — Auto MPG, spam detection, housing prices, telecom churn — and the intervals look sensible. And the computation time? We’re talking milliseconds. Lu, I bet you’ve got thoughts on why this matters for the broader research community.
Lu: Oh, absolutely, Jane. This is one of those papers that makes you wonder why nobody did it sooner. The out-of-bag estimate has been around since Breiman’s original random forest paper in two thousand one. But treating the whole training process as an augmented dataset and bootstrapping that — it’s elegant. It turns something that was a point estimate into a full inferential tool.
Tom: And Meng, from an engineering standpoint — is this something you’d actually use in production?
Meng: Honestly, Tom, yeah. When I’m shipping a model, I want to tell stakeholders not just “here’s the accuracy,” but “here’s the accuracy and here’s how much it might vary.” This gives me that without retraining a bunch of models. That’s a real win for deployment workflows.
Jane: And it’s all in an open-source R library called rangerror, built on top of the ranger package. So it’s not just theory — you can actually use it today.
Tom: Alright, so we’ve got the big picture. But how does the bootstrapping actually work under the hood? That’s what we’re digging into next.
Summary: Tom: So we’re back with “Confidence intervals for the random forest generalization error,” and Jane, I want to get into the mechanics now. How does this bootstrap trick actually work?
Jane: Okay, so picture this. When you train a random forest, each tree is grown on a bootstrap sample — a random sample of your training data, drawn with replacement. That means for each data point, about thirty-seven percent of the trees never saw it during training. Those trees are called “out-of-bag” for that point.
Tom: And the out-of-bag error is just the average prediction error computed using only those out-of-bag trees for each point. It’s a clever way to get a test error without holding out data.
Jane: Right. Now, the author’s insight is that you can collect all the pieces — the original data, which trees were out-of-bag for each point, and the predictions each tree made for each point — into what he calls an “augmented training sample.” Then you bootstrap that augmented sample. You resample the data points with replacement, recompute the out-of-bag error on that resampled set, and repeat thousands of times.
Tom: And those repeated out-of-bag errors give you the distribution, so you can read off the 2 point 5th and 97 point 5th percentiles for a ninety-five percent interval. It’s like a bootstrap within a bootstrap, but the inner bootstrap is already done — the forest is already trained.
Jane: Exactly. And that’s why it’s so cheap. You’re not retraining any trees. You’re just resampling the precomputed per-observation errors. The paper shows the whole thing runs in milliseconds, even for datasets with thousands of points.
Lu: What I find really clever is that this approach is automatically invariant to monotonic transformations. So if you’re predicting house prices and you want your confidence interval in dollars rather than squared dollars, you just transform the endpoints. The paper does exactly that with the Ames housing dataset — they report intervals in US dollars, which is way more interpretable.
Meng: And for classification, the intervals are guaranteed to stay within one, since the error is a proportion. That’s not true if you tried to use a normal approximation. So you avoid those silly intervals that go negative.
Tom: Meng, you’re an engineer — do you care about the coverage properties? The paper shows that the simulated coverage is pretty close to nominal, right?
Meng: Yeah, Tom, that’s the key thing. For the Friedman regression dataset with five hundred samples, a nominal ninety-five percent interval had actual coverage around ninety-two percent. For one thousand samples, it was ninety-four point three percent. It’s not perfect, but it’s close, and it gets better with more data. That’s what I need to trust it.
Jane: And the width shrinks as sample size grows, which is exactly what you want. The paper even measures the shrinking rate — it’s roughly n to the minus zero point seven six for the regression case, which is faster than the usual square-root rate.
Tom: So the intervals get tighter faster than you’d expect. That’s a nice bonus. But what’s the catch? There’s always a catch.
Jane: Well, the catch is that this is a bootstrap approximation, so it’s not exact. And the coverage isn’t perfect, especially for small samples. But for a method that costs almost nothing, it’s remarkably good.
Tom: Alright, so we’ve got the method and the results. But what does this mean for how we actually use random forests in practice? That’s up next.
Improvements: Tom: We’re back with “Confidence intervals for the random forest generalization error,” and I want to talk about what this actually improves. Jane, what’s the old way of getting a confidence interval for model error?
Jane: The old way is painful, Tom. You’d split your data into training and test sets, train the model on the training set, evaluate on the test set, and then repeat that whole process — maybe with cross-validation — to get multiple error estimates. Then you’d compute a standard deviation and build an interval from that.
Tom: And the problem with that?
Jane: You’re retraining the model dozens or hundreds of times. That’s computationally expensive, especially for big datasets. And you’re also throwing away data every time you hold out a chunk for testing.
Lu: The beauty of this paper is that it eliminates both problems. You train the forest once, on all your data. The out-of-bag mechanism already gives you a test-like error for every point. And then the bootstrap gives you the distribution. No retraining, no data splitting.
Meng: And that’s a real improvement for production. When I’m deploying a model, I don’t want to spend hours doing cross-validation just to get an error bar. This gives me an interval in milliseconds. That changes how often I can check model quality.
Tom: So it’s not just a theoretical nicety — it’s a practical tool.
Jane: Exactly. And the paper shows it works on real data. For the Spam dataset with four thousand six hundred one emails, the ninety-five percent interval for misclassification error is roughly four point one percent to five point four percent. For the telecom churn data, it’s about three point five percent to four point eight percent. Those are tight intervals, and they came almost for free.
Lu: What I find exciting is the potential for extension. The author uses a percentile bootstrap, but you could imagine using bias-corrected or accelerated methods to improve coverage further. The framework is flexible.
Meng: And you could also apply this to other bagging-based models, not just random forests. Gradient boosting doesn’t have the out-of-bag structure, but any ensemble that uses bootstrap sampling could benefit from this idea.
Tom: So the improvements are both immediate — you get intervals without extra training — and forward-looking — the framework could generalize.
Jane: Right. And the author also provides the rangerror R library, so you don’t have to implement it yourself. You just plug in your ranger model and get your intervals.
Tom: Meng, you mentioned production — how would you actually use this in a real system?
Meng: I’d use it as a monitoring tool. Every time I retrain a model, I’d compute the interval and log it. If the interval widens significantly over time, that’s a signal that the data distribution is shifting. It’s a cheap early-warning system.
Lu: That’s a great use case. And it also helps with model comparison — if two models have overlapping intervals, you know they’re not statistically different. If they don’t overlap, you have evidence one is better.
Tom: So this paper doesn’t just give you a number — it gives you a whole toolkit for making better decisions. And that’s what we’re wrapping up with next.
Conclusion: Tom: Alright, we’re at the end of our discussion on “Confidence intervals for the random forest generalization error.” Jane, give us the final summary.
Jane: So, the paper shows that the out-of-bag error you already get from training a random forest can be turned into a full confidence interval — just by bootstrapping the augmented training sample. No retraining, no data splitting, and it runs in milliseconds.
Tom: And the coverage is solid — close to nominal levels in simulations, and it improves with sample size. The intervals also shrink at a good rate, sometimes even faster than the usual square-root rule.
Lu: The real impact is that this makes uncertainty quantification accessible. You don’t need a PhD in statistics to get an error bar on your model. You just call a function.
Meng: And from an engineering perspective, it’s a practical tool for monitoring and model comparison. It gives you a cheap way to know if your model is degrading or if two models are actually different.
Tom: Lalam, you’ve been quiet — what’s your take on the bigger picture?
Lalam: I think this paper is part of a broader shift toward making machine learning more trustworthy. When every model comes with a confidence interval, it becomes easier to communicate uncertainty to non-experts — doctors, policymakers, business leaders. That builds trust in AI systems. And the fact that it’s open source means it can spread quickly.
Jane: And that’s the thing — this isn’t just a paper for academics. It’s a paper for anyone who builds models and wants to know how much to trust them.
Tom: Well said, Jane. We’ll be saying goodbye to “Confidence intervals for the random forest generalization error” — but we’re taking the idea with us. Next up, we’ve got another paper that’s been making waves, and I can’t wait to dig into it.
Jane: Thanks for listening, everyone. See you on the next episode.
Tom: Take care, and keep building.
Paulo C. Marques F
Insper Institute of Education and Research
stat.ML, cs.LG
Submitted: 2022-03-11
Updated: 2026-08-18
Comments: 10 pages
Code: https://github.com/aulocmarquesf/rangerror
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 69/100
Key concepts
- Out-of-bag error
- When training a random forest, certain data points are not used in the bootstrap sample for specific trees. The out-of-bag error is an estimate of how well the model will perform on new data, calculated using only these unused trees to provide a free measure of generalization.
- Confidence Interval
- A confidence interval provides a range (e.g., five percent plus or minus X) around a single point estimate, showing the level of certainty in the true value. The paper allows this to be calculated for model error by bootstrapping the augmented training sample.
- Bootstrapping
- Bootstrapping involves resampling data points with replacement and repeating calculations many times. This method is used here to generate a distribution of errors, allowing readers to derive specific percentiles (like the 2.5th or 97.5th) to define the confidence interval.
Terminology
Summary
Summary
This paper introduces a computationally inexpensive method for constructing confidence intervals for the generalization error of random forests, leveraging only the byproducts of the standard training process—specifically, the out-of-bag (OOB) estimates and the individual tree predictions on the training data—thereby avoiding any data splitting or model retraining.
The authors begin by framing the problem: How confident can we be in the generalization capacity of a predictive model?
They note that common tools like random train/test splits and cross-validation provide point estimates but raise the question of quantifying confidence without excessive computational cost. Random forests, through their bagging mechanism, offer a nearly free point estimate of the generalization error via the out-of-bag approach, since each training data point is not used (stays 'out-of-bag') when growing approximately 36.8% of the trees in the forest.
The core idea is to treat the original training data augmented with OOB bookkeeping and per-tree predictions as an augmented training sample,
which is then bootstrapped to produce confidence intervals. Formally, for a random forest ψ̂n, the generalization error is defined as γn = E[(Y − ψ̂n(X))2] in regression and γn = E[I(Y ≠ ψ̂n(X))] in classification. The OOB estimate is γ̂n, computed by averaging per-observation prediction errors using only trees for which each observation was out-of-bag.
The proposed method constructs the augmented training sample An = (xi, yi, Oi, ŷij Bj=1) ni=1, where Oi indexes the trees for which observation i was out-of-bag and ŷij = ψ̂(j)(xi) is the prediction of tree j for observation i. Then, for m = 1,..., M, the procedure uniformly samples with replacement n points from An to create An*(m), computes the corresponding OOB estimate γ̂n*(m), and finally obtains an approximate level 1−α confidence interval from the empirical α/2 and 1−α/2 percentiles of γ̂n*(1),..., γ̂n*(M). The computation is efficient because each γ̂n*(m) can be computed by directly resampling the per-observation errors γ̂(i) (squared error for regression, 0/1 loss for classification). The pseudocode is given in Algorithm 1.
The authors highlight two advantages of this percentile-based bootstrap approach over a central limit theorem approach: (1) it is automatically invariant to monotonic transformations of the quantity of interest (useful for reporting intervals in natural units, e.g., miles per gallon or US dollars), and (2) in classification, it guarantees that confidence intervals only contain valid values of the generalization error (which lies in [0,1]).
Simulation studies are used to assess effective coverage. For regression, they simulate from the Friedman process (ten independent predictors uniform on [0,1], response Y = 10 sin(πX1X2) + 20(X3 − 1/2)2 + 10X4 + 5X5 + ε with ε N(0,1), and the last five predictors being noise). For classification, they use a modified Gaussian spheres example (twenty standard normal predictors, response based on the sum of squares of the first ten predictors, with a 5% random flip of the sign). For both, they train random forests with B = 103 trees, use test samples of size ntst = 105 to approximate the true generalization error, and replicate the whole procedure N = 103 times. The results in Table 1 show that the simulated coverages are close to the corresponding nominal confidence levels
for training sample sizes n = 500 and n = 1000. The average widths of the confidence intervals shrink at rates approximately O(n−0.498) for the Gaussian spheres problem and O(n−0.764) for the Friedman process.
The method is then applied to four real datasets. For the Auto MPG dataset (n = 392, regression, predicting miles per gallon), the 95% confidence interval for the generalization error is [2.35, 3.09] (in miles per gallon). For the Spambase dataset (n = 4601, classification, spam detection), the 95% interval is [0.0411, 0.0537]. For the Ames housing dataset (n = 2930, regression, predicting sale price in US dollars), the 95% interval is [23,657.64, 27,788.31] (in US dollars). For the Telecom churn dataset (n = 3150, classification, predicting customer churn), the 95% interval is [0.0346, 0.0476]. These are reported in Tables 2 and 3.
Finally, the paper reports running times for Algorithm 1 on the four real datasets, based on 100 replications with M = 103 bootstrap replications. The median times are 68.8 ms for Auto MPG, 1008.7 ms for Spam, 580.9 ms for Ames housing, and 540.5 ms for Telecom churn, demonstrating the low computational cost of the procedure. The authors also point to an open source R library, rangerror, available at http://github.com/paulocmarquesf/rangerror, which implements the described procedures based on the ranger random forest library.
Improvements for AI systems
Based on the paper, here are the specific improvements I can implement in an AI system:
Current limitation: Most random forest implementations (e.g., scikit-learn, ranger) return only a point estimate of generalization error (the OOB error) without any uncertainty quantification.
Improvement: I will modify the random forest training pipeline to automatically compute and return confidence intervals for the generalization error using the bootstrapped augmented training sample method described in Algorithm 1.
What the improved system can do:
-
After training a random forest on any dataset, it outputs not just
oob error = 0.045but also95% CI = [0.0411, 0.0537](as in the Spam dataset example) -
This works for both regression (MSE) and classification (misclassification rate) without any additional data splitting or retraining
-
The computation is nearly free: it only requires resampling the already-computed per-observation OOB errors, adding 1 second for 4,600 samples with 1,000 bootstrap replications
Current limitation: When comparing two random forests (e.g., different hyperparameters), practitioners often pick the one with lower OOB error, ignoring that the difference may be within noise.
Current limitation: Users often train forests with a fixed number of trees (e.g., 500) without knowing if more trees would meaningfully reduce generalization error.
Current limitation: AI systems often report a single accuracy/MSE number in model cards or deployment logs, which can mislead stakeholders about true performance uncertainty.
Current limitation: Standard confidence intervals for regression errors assume normality (e.g., using t-distributions), which fails for skewed or heavy-tailed prediction errors common in real-world data (e.g., house prices, insurance claims).
Current limitation: Most systems only report overall performance, but fairness audits require per-group error rates (e.g., by age, gender, region).
Current limitation: Once deployed, models are often monitored using rolling accuracy, but there is no way to know if a drop is statistically significant without holding out a large labeled sample.
Current limitation: Grid search or Bayesian optimization for mtry, min node size, etc., uses point estimates of OOB error, wasting computation on hyperparameters that are statistically indistinguishable.
Capability Before After
Performance reporting Single OOB error number OOB error + 95% CI
Model comparison A is better than B
"A is better than B (p<0.05) or
No significant difference"
Training time Fixed number of trees Adaptive stopping based on CI stability
Fairness audit Point estimates per group Group-specific CIs with overlap testing
Production monitoring Manual threshold setting Automated drift detection using training CI
Hyperparameter tuning Compare point estimates Compare CIs to avoid overfitting to noise
These improvements are directly implementable, computationally inexpensive (adding <1% overhead to training), and provide statistically rigorous uncertainty quantification that can prevent costly deployment mistakes.
Abstract
We show that the byproducts of the standard training process of a random forest yield not only the well known and almost computationally free out-of-bag point estimate of the model generalization error, but also give a direct path to compute confidence intervals for the generalization error which avoids processes of data splitting and model retraining. Besides the low computational cost involved in their construction, these confidence intervals are shown through simulations to have good coverage and appropriate shrinking rate of their width in terms of the training sample size.
Related papers
- Behavior of prediction performance metrics with rare events
- Optimal Estimation of Generic Dynamics by Path-Dependent Neural Jump ODEs
- A Posterior-Dynamics Framework for Imaging Inverse Problems with Pretrained Diffusion Priors
- One Permutation Is All You Need: Fast, Deterministic Feature Importance and Model Stress-Testing
- Online Conformal Prediction for Non-Exchangeable Panel Data
- Deep Time-Series Forecasting in 10 Years: A Survey