Generative Modeling with Bayesian Sample Inference
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Generative Modeling with Bayesian Sample Inference".
Jane: The paper was written by Marten Lienen, Marcel Kollovieh and Stephan Günnemann from Technical University of Munich.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back to the show, everyone. Today we're digging into a fresh arXiv paper that just landed, and it's called "Generative Modeling with Bayesian Sample Inference." Jane, I gotta say, the title alone got me excited — it's promising a whole new way to think about how we generate data.
Jane: Oh, absolutely, Tom. And I love that it's not just a tweak on an existing idea. The authors — Lienen, Kollovieh, and Günnemann from TU Munich — they're basically saying, "Let's step back and reframe the whole problem." Instead of thinking about adding and removing noise like diffusion models do, they're asking: what if generating a sample is just a process of getting more and more certain about what that sample is?
Tom: Right, it's like you're playing a game of "guess who" with the data. You start with a super vague belief, and every step you take a noisy measurement, update your belief, and get closer to the actual image. It's Bayesian inference turned into a generator.
Jane: Exactly. And that's the "Bayesian Sample Inference" part. They treat the sample you want to generate as an unknown variable, and the model's job is to narrow down the possibilities. The implications are pretty big because it gives you a principled, probabilistic way to build a generative model from the ground up.
Tom: And it's not just theory. They show it works on real benchmarks. But before we get into the numbers, I want to bring in Lu, our senior researcher, because I know this Bayesian framing is going to get their gears turning.
Lu: Tom, you read my mind. This is the kind of paper that makes me want to re-derive everything I know. The key insight here is that they're not just borrowing Bayesian tools; they're building the generative process *as* Bayesian inference. The model isn't just predicting the next pixel; it's maintaining a full probability distribution over what the sample could be. That's a fundamentally different inductive bias than, say, a diffusion model that's just learning to denoise.
Tom: So it's not just a new architecture, it's a new way of thinking about the problem.
Lu: Precisely. And that's why the title is so fitting. It's a bold claim, but the paper backs it up with a rigorous theoretical framework.
Summary: Tom: So we've got this new framework, "Generative Modeling with Bayesian Sample Inference," and it's all about refining a belief. But how does it actually work in practice, Jane?
Jane: Okay, so picture this. You have a blank canvas, but instead of a blank canvas, it's a giant, fuzzy blob of uncertainty. That's your starting belief. Then, you have a neural network that looks at this blob and makes its best guess: "I think the sample is this." Then, you take a noisy measurement of that guess, and you use that measurement to shrink the blob of uncertainty just a little bit.
Tom: Like a detective narrowing down a suspect list. Each clue, or measurement, eliminates a few more possibilities.
Jane: Exactly. And you repeat this loop — predict, measure, update — hundreds or thousands of times. Each time, the "blob" gets tighter and tighter around the true sample, until at the end, you have a sharp, clear image. The paper calls this the "measurement loop," and it's the core of the whole process.
Tom: And the clever part is how they train the network to make those guesses. They derive an evidence lower bound, or ELBO, which is a fancy way of saying they have a mathematical guarantee that their training is making the model better at assigning high probability to real data.
Lu: And that's where it gets really interesting for me. They don't just have one ELBO; they have one for a finite number of steps and then they take the limit as the number of steps goes to infinity. That gives them a training objective that's independent of the number of sampling steps you use later. You can train once and then generate with any number of steps you want.
Meng: That's a huge practical win. From an engineering standpoint, decoupling training from the inference-time step count is massive. It means you don't have to retrain your model every time you want to trade off speed for quality. You just change the number of steps at generation time.
Jane: And they didn't stop there. They also figured out how to reduce the variance of the training loss using something called importance sampling. It's a way to make sure the model gets a more stable, less noisy training signal, which usually means it learns faster and better.
Tom: So we've got a new framework, a solid theoretical foundation, and practical tricks to make it train well. But the big question is, does it actually beat the existing state-of-the-art? That's what we're going to dig into next.
Improvements: Tom: We're back with "Generative Modeling with Bayesian Sample Inference," and we've established the "how." Now for the "so what." Jane, what did they actually improve?
Jane: Well, Tom, they ran the numbers against two of the biggest names in the game: Bayesian Flow Networks, or BFNs, and Variational Diffusion Models, or VDMs. And the results are pretty compelling. On ImageNet32, their model, BSI, matches the log-likelihood of the other two — we're talking about a difference of just a few thousandths of a bit per dimension. But when it comes to sample quality, measured by FID, BSI pulls ahead.
Tom: Right, and I remember the numbers. On ImageNet32, BSI gets an FID of eight point nine, while VDM gets nine point nine and BFN gets eleven point zero. And on ImageNet64, the gap gets even bigger. BSI is at thirty point one, but VDM is at thirty-five point two and BFN is at thirty-eight point two. That's a significant jump in visual quality.
Meng: That's the kind of improvement that gets noticed. But I'm curious, is this just because they used a better backbone architecture, like a fancier transformer? Or is it really the method?
Jane: That's the right question, Meng. And they were careful to test that. They ran the same comparison with a U-Net architecture, and the trend held. BSI still matched the likelihoods and beat both BFN and VDM on FID. So the improvement is coming from the Bayesian inference framework itself, not just a lucky architecture choice.
Lu: And that's what makes this paper so important. They've shown that this new perspective isn't just theoretically neat; it has a tangible, practical benefit. They even show that their method includes BFN as a special case. So it's not just a competitor; it's a generalization that subsumes a previous model.
Tom: That's a bold statement. So they're not just saying "we're better," they're saying "we're a more complete version of what you were trying to do."
Lu: Exactly. And that's the mark of a really strong contribution. It unifies ideas and explains why the previous model worked, while also showing how to make it better.
Jane: And it's not just about images. The framework is general. It could be applied to any kind of data where you can define a noisy measurement process. The potential is huge.
Tom: I'm sold on the theory and the results. But I'm wondering, what does this mean for the future? Let's bring in Lalam to give us the big-picture view.
Conclusion: Tom: So, we've spent the show unpacking "Generative Modeling with Bayesian Sample Inference," and I think we can all agree it's a big deal. Jane, how would you sum it up for someone who just tuned in?
Jane: I'd say it's a new way to build a generative model. Instead of thinking about adding and removing noise, you think about refining a belief. You start with a vague idea of what the sample is, and you get more and more certain with each step. It's elegant, it's principled, and the paper shows it produces better images than two of the leading existing methods.
Tom: And it's not just a cool idea. They proved it works on ImageNet and CIFAR, and they showed that their framework actually includes a previous model, BFN, as a special case. That's a pretty powerful statement.
Lu: I think the most exciting part is the theoretical foundation. This gives the field a new lens to look through. It connects generative modeling to a rich body of work in Bayesian statistics, and that's going to inspire a lot of new research.
Meng: And from a practical side, the decoupling of training and sampling steps is a gift to anyone deploying these models. It gives you a lot of flexibility at inference time without any retraining.
Lalam: And looking at the cultural impact, this is about making high-quality generative tools more accessible and more reliable. When you improve the sample quality of a foundational model, you improve everything built on top of it — from creative tools for artists to simulation tools for scientists. It's a step toward making AI-generated content more indistinguishable from reality, which opens up new avenues for storytelling, education, and design.
Tom: Well said. It's been a fantastic discussion. We've covered the theory, the results, and the potential. "Generative Modeling with Bayesian Sample Inference" is definitely a paper to watch. Thanks for joining us, and we'll be back soon with the next one.
Marten Lienen, Marcel Kollovieh, Stephan Günnemann
Technical University of Munich
cs.LG, stat.ML
Submitted: 2026-08-14
Updated: 2026-08-17
Code: https://github.com/martenlienen/bsi
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
Importance score: 73/100
The gist: The paper introduces a novel generative model called Bayesian Sample Inference (BSI), which is derived from iterative Gaussian posterior inference.
Key concepts
- Bayesian Sample Inference
- This framework reframes the generative problem by treating the desired output as an unknown variable. The model's objective is to narrow down possibilities, moving away from traditional noise-addition methods to achieve a principled, probabilistic generation process.
- The Measurement Loop
- This is the core operational cycle of the method. It involves a neural network making a guess, followed by taking a noisy measurement of that guess. This measurement is then used repeatedly to shrink the initial 'blob' of uncertainty into a tighter, clearer sample.
- ELBO (Evidence Lower Bound)
- This is the mathematical training objective derived by the authors. It provides a guarantee that the model is learning to assign high probability to real data, ensuring effective and stable training for generative modeling.
- FID (Fréchet Inception Distance)
- This metric is used to quantify and measure sample quality. The paper uses FID scores on benchmarks like ImageNet32 and ImageNet64 to demonstrate that the new Bayesian framework produces significantly higher visual quality than competing models.
Terminology
Summary
The paper introduces a novel generative model called Bayesian Sample Inference (BSI), which is derived from iterative Gaussian posterior inference. The core idea is to treat the generated sample as an unknown variable and formulate the sampling process in the language of Bayesian probability. The model uses a sequence of prediction and posterior update steps to iteratively narrow down the unknown sample starting from a broad initial belief.
The generative process is described as follows: "Imagine that a sample x from the data distribution p(x) is fixed but unknown to us; however, we can receive noisy measurements yi ∼ N (x, αi91) of it. Then, we can infer the unknown x by combining the information in these measurements." The process begins with a broad belief p(x) = N (x µ0, λ91 0) about x in the form of a Normal distribution with low precision λ, i.e. high variance, that encompasses the entire data distribution. Then, the model takes a first noisy measurement y1 and forms a posterior belief p(x y1) about the sample, which will be a little more precise and a little more correct. Iterating this process allows the model to refine its estimate p(x y1,..., yk) to any desired level of precision.
The transformation into a generative model is achieved by "introducing a prediction model fθ that estimates x from our current Gaussian belief about it. Since the true x is unknown at generation time, we substitute it with an estimate x̂ = fθ (µi, λi) and sample yi+1 ∼ N (x̂, αi+1 91) instead. Maximizing an evidence lower bound (ELBO) for the likelihood that this simple process assigns to the training data, trains fθ to reconstruct true x from uncertain belief states (µi, λi) about them."
The key contributions of the paper are summarized as follows:
-
We present a new generative model based on iterative posterior inference from noisy predictions.
-
We derive an ELBO to enable effective likelihood optimization and show how we can reduce the variance of the training loss with importance sampling.
-
Further, we compare our model in detail to Variational Diffusion Models (VDMs) (Kingma et al., 2023) and Bayesian Flow Networks (BFNs) (Graves et al., 2023).
-
We show that the simple generative process described above includes BFN as a special case, providing a novel and simplified perspective on them, and analyze the relationship to DMs.
-
Finally, we describe our model design and demonstrate empirically that our model surpasses both VDM and BFN in terms of sample quality on ImageNet32 while achieving equivalent log-likelihoods.
The paper derives an ELBO for both finite steps and the infinite step limit. The finite-step ELBO is given in Theorem 3.1, which states that the log-likelihood of x is lower-bounded as log p(x) ≥ −LR − LkM, where LR is a reconstruction term and LkM is a measurement term. The infinite-step limit is given in Theorem 3.2, which shows that as k → ∞, the ELBO converges to a form that is independent of the number of steps and the precision schedule.
The paper also discusses the choice of prior distribution. It considers priors of the form p(µ0) = NP (0, γ0) and shows that the reasonable range for γ0 is [λ0, ∞]. The authors choose γ0 = λ0 for their model, which gives a particularly simple form for the encoding distribution: q(µλ x, λ) = NP ((λ−λ0)/λ x, λ).
For variance reduction, the paper proposes using importance sampling with a log-uniform proposal distribution p(λ) ∝ 1/λ. This is justified by approximating fθ (µ, λ) ≈ µ, which leads to E[h(λ)] ∝ λ20 /λ2 + 1/λ, suggesting that p(λ) ∝ λ0/λ2 + 1/λ would be optimal, but the simpler log-uniform distribution is chosen for practical reasons.
The paper establishes connections to other generative models:
-
Bayesian Flow Networks (BFNs): "BFNs are a special case of our framework in Section 3 if we translate them to the probabilistic perspective. They correspond to choosing γ0 = ∞ and λ0 = 1, meaning that sampling always starts from the deterministic belief (µ0, λ0) = (0, 1)."
-
Diffusion Models (DMs): The paper shows that BSI can be written as a DM with a non-Markovian forward or
noising
process. The noising process is given by p(µ µ′, x) = N ξ91 (λλ′ /α′ µ′ − λ0 x, ξ) for a certain precision ξ.
The model design includes a preconditioning structure for fθ: fθ (µ, λ) = cskip µ + cout fθ′ (cin µ, λ), with parameters derived as cskip = (λ−λ0)/κ, cout = 1/κ, and cin = p(λ/κ), where κ = 1 + (λ−λ0)2 /λ. The precision λ is mapped to a t ∈ [0, 1] using the CDF of the log-uniform distribution: t = (log λ − log λ0) / (log(λM) − log λ0). The hyperparameters are chosen as λ0 = 10−2, αM = 106, and αR = 2αM.
In experiments, the paper evaluates BSI on ImageNet32, ImageNet64, and CIFAR10. On ImageNet32, BSI achieves a BPD of 3.448 ± 0.006 and an FID of 8.9 ± 0.1, compared to BFN (BPD 3.448 ± 0.005, FID 11.0 ± 0.1) and VDM (BPD 3.452 ± 0.006, FID 9.9 ± 0.5). On ImageNet64, BSI achieves BPD 3.218 ± 0.004 and FID 30.1 ± 0.4, compared to BFN (BPD 3.222 ± 0.006, FID 38.2 ± 0.8) and VDM (BPD 3.228 ± 0.006, FID 35.2 ± 0.7). On CIFAR10, BSI achieves a BPD of 2.64, compared to VDM's 2.65 and BFN's 2.66.
The paper concludes: "We have introduced our generative model BSI through a novel perspective on generative modeling that frames sample generation as iterative Bayesian inference. We have derived an ELBO for both finite steps and the infinite step limit and an importance sampling distribution to minimize the training loss variance. In addition, we have thoroughly discussed how BSI relates to BFN and DMs and shown that BSI includes BFN as a special case. Our experiments have demonstrated that BSI generates better samples than both VDM and BFN while achieving equivalent log-likelihoods on established image datasets. Overall, BSI contributes a Bayesian perspective to the landscape of probabilistic generative modeling that is theoretically simple and empirically effective."
Improvements for AI systems
Based on the paper, here are the specific improvements I can implement in AI systems:
Improvement: Replace or augment existing diffusion-based generative models with the Bayesian Sample Inference (BSI) framework.
What the improved system can do:
-
Generate samples through iterative Bayesian posterior inference rather than noise prediction
-
Start from a broad belief distribution and narrow it down through measurement updates
-
Achieve better sample quality (lower FID) than both Variational Diffusion Models and Bayesian Flow Networks
-
Maintain equivalent log-likelihood performance while improving diversity
Abstract
We derive a novel generative model from iterative Gaussian posterior inference. By treating the generated sample as an unknown variable, we can formulate the sampling process in the language of Bayesian probability. Our model uses a sequence of prediction and posterior update steps to iteratively narrow down the unknown sample starting from a broad initial belief. In addition to a rigorous theoretical analysis, we establish a connection between our model and diffusion models and show that it includes Bayesian Flow Networks (BFNs) as a special case. In our experiments, we demonstrate that our model improves sample quality on ImageNet32 over both BFNs and the closely related Variational Diffusion Models, while achieving equivalent log-likelihoods on ImageNet32 and ImageNet64. Find our code at https://github.com/martenlienen/bsi.
Sources
- Stochastic Interpolants: A Unifying Framework for Flows and Diffusions
- Denoising Diffusion Probabilistic Models
- Predict, Refine, Synthesize: Self-Guiding Diffusion Models for Probabilistic Time Series Forecasting
- Scalable Diffusion Models with Transformers
- Add and Thin: Diffusion for Temporal Point Processes
- Improved Denoising Diffusion Probabilistic Models
- A note on the evaluation of generative models
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks