Population Predictive Checks

arXiv:1908.00882 · stat.ME, cs.LG · Submitted 2026-08-14 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Population Predictive Checks".

Jane: The paper was written by Gemma E. Moran, David M. Blei and Rajesh Ranganath from Columbia University and New York University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title and Authors: Tom: Welcome back to the show, everyone. Today we are digging into a paper that just hit arXiv, and it’s called “Population Predictive Checks.” Jane, I have to say, the title alone got me excited because it sounds like it’s fixing something that’s been bugging statisticians for decades.

Jane: Oh, absolutely, Tom. And the authors are Gemma Moran from Columbia, David Blei from Columbia, and Rajesh Ranganath from NYU. These are heavy hitters in Bayesian modeling. Blei basically wrote the book on topic models, and Ranganath has done tons of work on variational inference. So when they team up to talk about model checking, you know it’s going to be good.

Tom: Right, and for our listeners who might not be deep in the weeds of Bayesian stats, let’s break down what “model checking” even means. You build a model to explain your data, but how do you know if the model is actually any good? That’s the whole game here.

Jane: Exactly. And the classic tool for this is called a posterior predictive check, or PPC for short. The idea is you fit your model to your data, then you generate new fake data from that fitted model, and you compare the fake data to your real data. If they look similar, your model passes. If they don’t, your model fails.

Tom: And that sounds reasonable, but the paper points out a sneaky problem. You’re using the same data twice. Once to fit the model, and once to check the model. And that’s like grading your own homework after you’ve already seen the answers.

Jane: That’s the “double use of the data” problem, and it makes the check overconfident. The model can look great even when it’s actually terrible, because it’s already memorized the data you’re testing it on. This paper proposes a fix called the population predictive check, or pop-pc.

Tom: And the fix is elegant. Instead of checking your model against the same data you used to fit it, you hold out a fresh chunk of data, like a test set, and you compare your model’s fake data to that fresh data. That way, the model hasn’t seen it before, so it can’t cheat.

Jane: Right. And the authors prove that this new check is properly calibrated, meaning the p-values you get are actually trustworthy. The old PPC could give you a p-value of zero point five all the time, which sounds fine but tells you nothing. The pop-pc gives you a real distribution, so you can actually tell if something is wrong.

Tom: And they show this works in practice, too. They run it on regression models and on topic models for text data, and the pop-pc catches overfitting that the old PPC completely misses. That’s a huge deal for anyone who uses Bayesian models in real research.

Jane: It really is. And the implications go beyond just fixing a statistical bug. This changes how we trust our models. If you’re using a model to make decisions, you need to know when it’s lying to you. The pop-pc gives you a much better lie detector.

Tom: And that’s what we’re going to dig into next. We’ll talk about the actual method, how it works step by step, and why it’s such a big improvement. Stick around.

Summary of the Paper: Tom: So we’re back with “Population Predictive Checks,” and Jane, we’ve set the stage. Now let’s get into the nitty-gritty of how this thing actually works.

Jane: Sure. So the setup is you have your observed data, call it y-obs. You fit your Bayesian model, which gives you a posterior distribution over your parameters. Then you draw replicated data from that posterior, call it y-rep. That’s the same as the old PPC.

Tom: But here’s where it changes. In the old PPC, you compare y-rep to y-obs. In the pop-pc, you split your data. You use one chunk to fit the model, and you keep a separate chunk, call it y-new, that the model has never seen. Then you compare y-rep to y-new.

Jane: And the key insight is that y-new is a draw from the true population distribution, not from your model. So you’re asking, “Does my model generate data that looks like real data from the world, or does it just look like the data it was trained on?” That’s a much tougher test.

Tom: And the paper formalizes this with a p-value. You calculate the probability that the diagnostic statistic of your replicated data is greater than the diagnostic statistic of your new data. If that probability is very low, your model is producing data that’s weird compared to the real world.

Jane: Right. And they prove in Theorem one that these p-values are asymptotically uniform when the model is correct. That means if your model is actually right, you’ll get p-values spread evenly between zero and one. If your model is wrong, the p-values will cluster near zero, telling you to reject it.

Tom: And that’s the calibration property we talked about. The old PPC doesn’t have that. In fact, the paper shows that the old PPC can get stuck at zero point five no matter what, which is useless. It can’t tell you if your model is good or bad.

Jane: And they don’t just stop at theory. They run experiments. In the regression example, they show that when you have too little regularization, the model overfits. The pop-pc catches this and rejects the model. The old PPC just sits there at zero point five, completely blind.

Tom: And they also test it on latent Dirichlet allocation, which is a topic model for documents. They show that when you have too many topics, the model starts memorizing the data. Again, the pop-pc catches it, and the old PPC doesn’t.

Jane: One thing I really appreciate is that they address the criticism that you could just calibrate the old PPC post-hoc. There are methods by Robins and by Hjort that try to fix the calibration issue. But the paper shows that even if you calibrate the p-values, they still don’t have power. They can’t detect model misfit.

Tom: That’s a crucial point. Calibration alone doesn’t fix the double use problem. You need a fundamentally different approach, which is what the pop-pc provides. It’s not just a patch; it’s a redesign.

Jane: And that redesign is what we’re going to explore next. We’ll talk about the improvements it suggests for how we build and check models in practice.

Improvements Suggested by the Paper: Tom: Alright, we’re back with “Population Predictive Checks,” and Jane, we’ve covered the basics. Now let’s talk about what this paper actually changes for practitioners. What does it suggest we do differently?

Jane: The biggest improvement is that it gives you a reliable way to detect overfitting. And overfitting is the silent killer in machine learning. You train a model, it looks great on your training data, and then it falls apart on new data. The pop-pc catches that because it tests on held-out data.

Tom: And the paper shows this really clearly in the regression example. They have a ridge regression model with a prior variance parameter, c. When c is small, there’s a lot of regularization, and the model generalizes well. The pop-pc keeps the model. When c is large, there’s no regularization, and the model overfits. The pop-pc rejects it.

Jane: But the old PPC, and even the calibrated versions, they just sit at zero point five. They can’t tell the difference between a well-regularized model and an overfit mess. That’s a massive improvement in practical utility.

Tom: And it’s not just about regression. They apply it to topic models, which are used all over the place for text analysis. They show that when you increase the number of topics to the point where the model is just memorizing the vocabulary, the pop-pc rejects it. The old PPC doesn’t.

Jane: Another improvement is that the pop-pc is simple to implement. You just split your data, fit your model on one half, and check it on the other half. You can reuse the same posterior samples for multiple diagnostics. The paper mentions that the partial predictive check, which is another alternative, requires you to redo the calculation for each diagnostic function.

Tom: That’s a practical win. If you’re a researcher, you don’t want to write a new algorithm every time you want to try a different diagnostic. With the pop-pc, you fit once, and then you can check as many things as you want.

Jane: And the paper also discusses how this connects to other ideas like intrinsic Bayes factors and cross-validation. It’s not a totally new concept, but it formalizes it in a way that’s rigorous and gives you theoretical guarantees.

Tom: And those guarantees matter. The paper proves that the pop-pc is calibrated, which means you can trust the p-values. That’s a big deal for reproducibility in science. If you’re going to reject a model based on a p-value, you want to know that p-value actually means something.

Jane: Right. And they also show that the pop-pc has high power. It can detect model misspecification with probability one in the limit. That’s the kind of guarantee you want when you’re making decisions based on your models.

Tom: So the improvements are clear: better detection of overfitting, simpler implementation, and theoretical guarantees. But what does this mean for the broader world? That’s what we’re going to wrap up with next.

Conclusion: Tom: And we’re back for the final segment on “Population Predictive Checks.” Jane, we’ve covered the method, the theory, and the experiments. Let’s pull it all together.

Jane: So the takeaway is that this paper gives us a new tool for Bayesian model criticism that actually works. The old posterior predictive check had a fatal flaw: it used the data twice, making it overconfident and unable to detect overfitting. The population predictive check fixes that by using held-out data.

Tom: And the authors prove it’s calibrated, which means the p-values are trustworthy. They also show it has high power, so it can actually reject bad models. And they demonstrate it on real problems, like regression and topic modeling.

Jane: The implications are huge for anyone who uses Bayesian models in practice. Whether you’re analyzing clinical trials, building recommendation systems, or doing text analysis, you need to know when your model is lying to you. The pop-pc gives you that ability.

Tom: And it’s not just about fixing a bug. It’s about changing the workflow. Instead of fitting a model and hoping it’s good, you can now actively test it against fresh data. That’s a more rigorous, more scientific approach.

Jane: And the paper is also honest about its limitations. The calibration proof relies on asymptotic normality of the diagnostic, which doesn’t always hold. But they show empirically that it still works well for things like the chi-squared diagnostic and the log-likelihood in topic models.

Tom: So it’s not a silver bullet, but it’s a big step forward. And it opens up avenues for future research, like extending it to hierarchical models and checking individual components of a model.

Jane: Absolutely. And we should say goodbye to this paper. It’s been a great discussion, and I think this is one of those papers that will become a standard reference for model checking.

Tom: Agreed. “Population Predictive Checks” is a must-read for anyone serious about Bayesian modeling. Thanks to Gemma Moran, David Blei, and Rajesh Ranganath for this contribution.

Jane: And thanks to our listeners for tuning in. We’ll be back next time with another paper from arXiv. Until then, keep questioning your models.

Tom: And keep your data separate. See you all next time.

Gemma E. Moran, David M. Blei, Rajesh Ranganath

Columbia University · New York University

stat.ME, cs.LG

Submitted: 2026-08-14

Updated: 2026-08-17

DOI: 10.1093/jrsssb/qkad105

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 69/100

The gist: The paper introduces the population predictive check (pop-pc), a new method for Bayesian model criticism that addresses the "double use of the data" problem inherent in posterior predictive checks

Terminology

Summary

The paper introduces the population predictive check (pop-pc), a new method for Bayesian model criticism that addresses the double use of the data problem inherent in posterior predictive checks (ppcs).

The paper identifies a crucial issue with the standard posterior predictive check (ppc): it uses the data twice. The data are first used to construct the reference distribution of the diagnostic, the posterior predictive distribution... The data are then used again in the observed diagnostic. This leads to two consequences: the ppc p-value may be likely to retain an incorrect model; it may suffer from low power and the ppc p-value may be overconfident about the correct model; it is not calibrated.

The pop-pc is defined as follows: "Consider observed data y obs, its posterior predictive distribution p(y rep y obs), and a diagnostic statistic d (y). Suppose we have y new drawn from the population distribution of the data. As a p-value, the population predictive check is: p pop = p(d (y rep) ≥ d (y new) y obs, y new), where y rep ∼ p(y rep y obs)."

The premise is: If my model is good, then data drawn from the posterior predictive distribution will look like a draw from the true population. To implement, one splits the data into y = (y obs, y new), assuming y new is an independent draw from the population distribution.

The paper proves the pop-pc is calibrated: "Theorem 1: Assume Equation 16 holds and assume the regularity conditions detailed in Appendix A.1 of the Supplementary Material. Under the distribution f (y; theta 0), the pop-pc p-value can be written as: p pop (y) = 1 − Φ(Q) + o P (1), where o P (1) denotes a random variable converging to zero in probability, Φ is the standard normal cdf, and Q ∼ N (0, 1). Consequently, the pop-pc p-values are calibrated."

The proof applies to diagnostic functions d (y) that are asymptotically normal with asymptotic mean nu(theta) and asymptotic variance sigma 2 (theta) when the model is correct.

The paper demonstrates that post-hoc calibration of ppcs is insufficient: calibration by itself does not necessarily improve the power of the check. Specifically, "these calibration techniques cannot detect model misfit because for any random variable, we can calibrate it by using its cdf to transform it to a uniform distribution. If the original random variable does not detect model misfit, this transformation will not provide additional power to detect model misfit."

In a specific Gaussian mean example, the paper shows: the ppc p-value is degenerate at 0.5 regardless of whether the data is Gaussian or Cauchy and the calibrated ppc is uniform regardless of whether the data is Gaussian or Cauchy. In contrast, the pop-pc is uniform when the data is Gaussian, and concentrated around 0 when the data is Cauchy. That is, the pop-pc detects model misspecification while the ppc or calibrated ppc do not. The paper proves the pop-pc has asymptotic power of one–that is, pop-pc will reject the Gaussian model if the data is actually Cauchy with probability one.

Bayesian Ridge Regression: The paper shows the ppc p-values are constant for all values of the regularization parameter, c; that is, the ppc cannot detect model misfit for large values of c and "the two calibrated ppcs (Robins et al., 2000; Hjort et al., 2006) also cannot detect this model misfit. Meanwhile, the pop-pc retains the model for small values of c. For large values of c, the pop-pc rejects the model as it overfits to the observed data. The pop-pc p-values are approximately uniform for small values of c and for large values of c, the p-values concentrate around 0."

Topic Modeling (LDA): On synthetic data, the pop-pc p-values are approximately uniformly distributed while the ppc p-values are left-skewed, providing empirical evidence that the pop-pc is calibrated for the LDA log-likelihood diagnostic. On New York Times data, as the number of topics increases to K = 1, 000, the ppc p-value remains close to 0.5 while the pop-pc p-value detects this overfitting, and begins to reject the model past K = 200 topics.

The paper notes connections to prior predictive checks (Box, 1980), partial predictive checks (Bayarri and Berger, 2000), cross-validated checks (Gelfand et al., 1992), and intrinsic Bayes factors (Berger and Pericchi, 1996). It also acknowledges concurrent work by Li and Huggins (2022) on split predictive checks, noting This pop-pc is the same as the single spc of Li and Huggins (2022).

The pop-pc is simple to implement, and the inferred posterior can be used to check many different diagnostic functions, unlike the partial predictive check which can be difficult to calculate, and requires re-calculation for each diagnostic function.

Improvements for AI systems

Based on the scientific paper, here are specific improvements for AI systems:

  • Replace or supplement posterior predictive checks (ppc) in AI systems that use Bayesian inference (e.g., Bayesian neural networks, Gaussian processes, hierarchical models).

  • Split data into training and heldout sets for model evaluation: use one half for posterior inference, the other half for diagnostic checking.

  • Calculate pop-pc p-values using the formula: p(d(y rep) ≥ d(y new) y obs, y new), where y rep is drawn from the posterior predictive and y new is heldout data.

  • Flag overconfident models that ppc would incorrectly accept. The paper shows ppc p-values concentrate around 0.5 regardless of model fit, while pop-pc correctly rejects overfit models (e.g., when regularization is too weak).

  • Use pop-pc to tune hyperparameters like prior variance in Bayesian regression, or number of topics in LDA, by rejecting configurations where pop-pc p-values drop below significance thresholds.

  • Guarantee calibrated p-values under asymptotic normality of diagnostics (Theorem 1). This means AI systems can reliably interpret p-values: a pop-pc p-value of 0.04 means a 4% chance of false rejection when the model is correct.

  • Avoid the double use of data problem that plagues ppc—the pop-pc does not use the same data to both construct and evaluate the reference distribution.

For Bayesian Neural Networks:

  • Detect when a network is memorizing training data rather than learning generalizable patterns (as shown in the regression example with large prior variance).

  • Use pop-pc with diagnostics like χ2 or log-likelihood to validate uncertainty estimates.

For Topic Models (LDA):

  • Automatically determine the optimal number of topics by rejecting models where pop-pc p-values indicate overfitting (as shown in the New York Times example, where ppc failed to reject even at K=1000 topics while pop-pc rejected beyond K=200).

  • Validate that topic assignments generalize to heldout documents.

For Hierarchical/Grouped Data Models:

  • Check whether group-level parameters are well-calibrated by comparing posterior predictive draws to heldout groups.

  • Algorithm 1 provides a straightforward implementation: draw posterior samples, generate replicated data, compute diagnostic, and compare to heldout diagnostic.

  • Use with any diagnostic function (χ2, log-likelihood, mean, etc.)—the method is diagnostic-agnostic.

  • Computational cost is similar to ppc, requiring only an additional heldout set and diagnostic computation on it.

  • Do not rely solely on post-hoc calibration of ppc (e.g., Robins et al. 2000 or Hjort et al. 2006)—these produce uniform p-values but fail to detect model misspecification (as proven in Section 4.2 and shown in Figures 4-5).

  • Use pop-pc instead of partial predictive checks when computational simplicity is needed—pop-pc is easier to implement and does not require re-calculation for each diagnostic.

  • Higher statistical power to detect model misspecification compared to ppc and calibrated ppc.

  • Correct calibration ensures false rejection rates match significance levels.

  • Better model selection in practice, as demonstrated by the regression and LDA experiments where pop-pc correctly rejected overfit models that ppc accepted.

Abstract

Bayesian modeling helps applied researchers articulate assumptions about their data and develop models tailored for specific applications. Thanks to good methods for approximate posterior inference, researchers can now easily build, use, and revise complicated Bayesian models for large and rich data. These capabilities, however, bring into focus the problem of model criticism. Researchers need tools to diagnose the fitness of their models, to understand where they fall short, and to guide their revision. In this paper we develop a new method for Bayesian model criticism, the population predictive check (Pop-PC). Pop-PCs are built on posterior predictive checks (PPCs), a seminal method that checks a model by assessing the posterior predictive distribution on the observed data. However, PPCs use the data twice -- both to calculate the posterior predictive and to evaluate it -- which can lead to overconfident assessments of the quality of a model. Pop-PCs, in contrast, compare the posterior predictive distribution to a draw from the population distribution, a heldout dataset. This method blends Bayesian modeling with frequenting assessment. Unlike the PPC, we prove that the Pop-PC is properly calibrated. Empirically, we study Pop-PC on classical regression and a hierarchical model of text data.

Sources

Related papers