BAM! Bayesian Anything Model: a foundation model for generative computational imaging
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "BAM! Bayesian Anything Model".
Tom: Generative models are transforming Bayesian computational imaging, yet the field still lacks physics-aware foundation models.
Jane: First, who's behind it and why it matters.
Paper summary: Tom: Welcome back everyone! We're diving into some really interesting work today about generative models and how they're being applied to computational imaging. We’ve got a paper called "BAM! Bayesian Anything Model: a foundation model for generative computational imaging" that sounds like it’s tackling a big challenge in this area.
Jane: It does sound substantial, Tom. This paper introduces something called the BAM (Bayesian Anything Model), and the idea is that current generative models aren't fully physics-aware yet, which is a key problem they are trying to solve.
Lu: I'm really intrigued by the idea of moving towards physics-aware foundation models; it suggests we can get past the bias you mentioned when using zero-shot approximate likelihood guidance in large foundation image models.
Meng: From an engineering standpoint, if these models can generalize robustly to unseen data and tasks with minimal finetuning, that opens up a lot of possibilities for deploying imaging solutions quickly.
Lalam: I think the potential here is huge for how we build and interact with vision-based systems; imagine culture shifts when the underlying models become this flexible.
Tom: Exactly. So, what's BAM actually proposing in terms of its core thesis? What is its main claim about what it can do?
Jane: The paper claims that BAM introduces a lightweight foundation model designed for few-step, physics-aware posterior sampling that generalizes robustly to unseen data and tasks with zero-shot or minimal finetuning.
Lu: It seems the core idea is upgrading an existing operator-conditioned Reconstruct Anything Model backbone into a conditional flow map, allowing instrument physics to be specified at inference time instead of being fixed during training.
Meng: So, it’s not just a new model architecture; it’s fundamentally changing how we handle the relationship between measurement and image reconstruction by making the physics dynamic during use.
Lalam: That dynamic specification sounds incredibly powerful because it decouples the instrument knowledge from the model itself, which is a big step for flexibility.
Tom: Right, so instead of locking in physics during training, BAM lets you specify those conditions when you actually need to sample something new. It’s about making the model adaptable on demand.
Jane: Precisely. They are focusing on imaging problems where we have an unknown image x and a measurement y related by y = Ax⋆ + σyw, where A and sigma y are known at inference time but the posterior p(x y, A, σy) is what we want to sample from.
Lu: And they handle varying scales by conditioning on a "rescaled measurement," defined through the stochastic interpolant yσ = αsAx + ςσw, which helps it generalize across different operators and noise levels encountered in practice.
Paper summary: Meng: That mechanism for handling scale variation is interesting from a practical standpoint because real-world data rarely fits a perfectly uniform setup. It suggests a more robust way to handle messy experimental conditions.
Lalam: When you combine that with the few-step sampling, it implies that getting high-quality samples doesn't have to take an enormous computational investment every single time we run an inference.
Tom: That leads us perfectly into the methodology—how does this model actually learn this flow map in the first place? We need to understand how they train BAM!
Jane: The training involves two main flow map training objectives. First, on the diagonal where t equals s, they fit the velocity to the interpolant slope by minimizing a loss function Lb(θ) involving E
vθ(xt, t, t, yσ, A) − (z − xzero)two/two: .
Lu: That objective seems designed to make sure the model’s movement along the flow map aligns correctly with the expected trajectory dictated by the measurement structure.
Meng: Minimizing that loss ensures that when we sample a point on that diagonal, it respects the relationship between time and the underlying data points in a structured way.
Lalam: It’s about enforcing consistency in how quickly or slowly the model evolves through these states, which is crucial for reliable sampling.
Tom: And then there's the off-diagonal part, where s is less than t, governed by the Lagrangian condition LLSD(θ), which looks at backwards jumps from t to s using E
∂sXθt,s − sg vθ−(xˆt,s, s, s, yσ, A)two/two: .
Jane: That second loss targets the backwards jumps between states and is governed by the Lagrangian condition of section two of the paper. It’s designed to regulate those transitions properly.
Lu: The objective L(θ) also includes auxiliary losses for perceptual quality and contrast bias, specifically a squeeze-based perceptual loss weighted by g(s) = exp(−4s), which is only weighted when s is small.
Meng: That weighting scheme on the perceptual loss tells us that the model cares most about sharp details or low-level structure when it’s making fine adjustments to the reconstruction.
Lalam: And they also use a contrast bias loss, lctr, which compares intensity histograms and is set to zero after pre-training to focus on structural fidelity first.
Tom: It sounds like they've balanced structural accuracy with perceptual quality very carefully during the training process itself. So, what are the main contributions we need to take away from this entire BAM! Bayesian Anything Model paper?
Paper summary: Jane: The key contributions are three things: proposing a flow-map training paradigm, developing a lightweight foundation model built on the 36M-parameter RAM backbone that learns this map across operators and noise levels, and achieving state-of-the-art quality at a fraction of the cost.
Lu: I think the flow map training paradigm is really significant because it upgrades an unfolded reconstruction network into this structure via Lagrangian self-distillation on a normalized stochastic interpolant of the measurement.
Meng: The lightweight nature, built on that 36M-parameter backbone, is what makes this practical for deployment rather than just theoretical work in a lab.
Lalam: This foundation model aspect means we don't have to start from scratch for every new imaging task; we can leverage this learned structure across different datasets and operators.
Tom: And the performance metrics are striking—outperforming specialized models and leading zero-shot methods in sample quality in just three steps, all at a fraction of the computational cost. That’s what really stands out to me from the experimental results.
Jane: They showed this across several linear inverse problems on FFHQ, AFHQ, LSUN, DIV2K and even the Köhler camera-shake benchmark. BAM also shows robustness under post-training quantization at INT8 with a speedup of three point five one times with only a moderate accuracy loss.
Lu: The fact that it provides a spatial uncertainty map for sparse-view CT problems at no extra training cost is an interesting addition, showing it has capabilities beyond just producing the final image samples.
Meng: That uncertainty map feature is something I’d look at closely for real-world applications where reliability and knowing what the model doesn't know are important factors in a decision.
Lalam: If this technology can be deployed widely because of its efficiency, it really means that complex computational imaging capabilities become accessible to much more people and systems.
Tom: So, we’ve seen how BAM! Bayesian Anything Model moves beyond just being another generative model by incorporating physics-aware flow maps for robust sampling. Now we need to think about what this means for the future of the field.
Jane: The paper suggests that this approach can provide a way to move away from models that rely on fixed priors or approximate likelihood guidance, which have introduced bias and cost issues.
Lu: From a theoretical perspective, it suggests a path toward building foundation models that inherently understand the underlying physical process rather than just memorizing data distributions.
Meng: For practical deployment, this means we might see imaging systems that can adapt to new sensor configurations or noise characteristics much more easily without requiring extensive retraining.
Lalam: It implies a future where generative AI in vision tasks isn't just about generating pretty pictures, but about generating physically plausible and robust solutions that work reliably in varied conditions.
Conclusion: Tom: So we've been deep in the weeds on this "BAM! Bayesian Anything Model" paper, and now it's time to wrap up with some big-picture thinking about what all this means for us.
Jane: I think that really captures the essence of the work—taking a model that generates images and making it much more grounded in how real-world physics works during the sampling process.
Lu: From a theoretical standpoint, BAM moves us away from purely statistical approximations toward something that respects the underlying physical equations, which is fascinating for developing more reliable generative models.
Meng: I'm thinking about how this will translate into actual hardware deployment; if it can handle varying noise and operators without needing a whole new training pipeline for every single sensor setup, that’s a huge win for practical engineering.
Lalam: It really shows us how AI can improve culture by enabling more sophisticated visual analysis tools that aren't brittle when they encounter unexpected real-world conditions.
Tom: Exactly! The authors, who are leading researchers in the imaging space, have built something quite clever here—a foundation model that learns to sample images efficiently without needing tons of specialized training for each specific task.
Jane: And the title itself, "Bayesian Anything Model," hints at its power because it suggests we can generate high-quality outputs across a whole range of possibilities with minimal effort.
Lu: The implication is that instead of building separate models for every new imaging problem or noise level, we can leverage this one foundation model and just give it the right conditions at inference time.
Meng: So, when you think about the impact on industry, it's about making complex imaging capabilities accessible to more people because the computational cost drops significantly.
Lalam: For me, I see this as a way to democratize high-quality visual synthesis; it means we can build applications that rely on these models being reliable in diverse environments.
Tom: That’s what I love—moving from theoretical potential to practical utility, and it looks like BAM is making real headway there with its efficiency and generalization.
Jane: It certainly does, but we still have to consider the limitations they laid out regarding the types of noise it handles and the need for specific instrument models in some scenarios.
Lu: Those are fair caveats; they're not claiming perfection, which keeps it grounded in reality rather than overpromising capabilities.
Meng: That makes sense; we need to know exactly where this model stops being useful so we can build the necessary safeguards into our systems.
Lalam: Knowing those boundaries helps us focus on where this technology will have the most meaningful cultural impact moving forward.
Alessio Spagnoletti, Charlesquin Kemajou Mbakam, Jonathan Spence, Andrés Almansa, Marcelo Pereyra
Laboratoire MAP5, UMR 8145, Université Paris Cité, CNRS 2Heriot-Watt University, School of Mathematical and Computer Sciences & Maxwell Institute for Mathematical Sciences
cs.CV, stat.ML
Submitted: 2026-09-30
Updated: 2026-09-30
Code: https://github.com/matthieutrs/ram3https:
Project page: https://bayesian-anything-model.github.io
Importance score: 92/100
The gist: Generative models are transforming Bayesian computational imaging, yet the field still lacks physics-aware foundation models.
Key concepts
- Flow Map Training Paradigm
- BAM transforms the reconstruction network into a 'flow map' by using Lagrangian self-distillation on a normalized measurement. This specific training recipe teaches the model how to move between different states in an image reconstruction process, allowing it to generalize across varying operators and noise levels effectively.
- Operator-Conditioned Reconstruct Anything Model (RAM)
- BAM uses the 36M-parameter RAM backbone as its foundation. This backbone is pre-trained jointly on many images and forward operators. BAM modifies this backbone into a conditional flow map, meaning instrument physics can be specified at inference time rather than being fixed during initial training.
- Rescaled Measurement
- To handle different scales in imaging problems, BAM conditions on a 'rescaled measurement.' This is defined using a stochastic interpolant that incorporates schedules ($\alpha_s$ and $\varsigma_s$) designed to satisfy the relationship between noise level and scale, ensuring the model works across diverse practical scenarios.
Terminology
Summary
Generative models are transforming Bayesian computational imaging, yet the field still lacks physics-aware foundation models. BAM introduces a lightweight foundation model for few-step, physics-aware posterior sampling that generalizes robustly to unseen data and tasks, zero-shot or with minimal finetuning.
How it works
BAM upgrades the operator-conditioned Reconstruct Anything Model (RAM) backbone into a conditional flow map, allowing instrument physics to be specified at inference time rather than fixed during training. It is built on the 36M-parameter RAM backbone and is pre-trained jointly on large image corpora and libraries of forward operators. The network draws posterior samples in a few steps, with no likelihood approximation and no guidance weights to tune.
The model is designed for imaging problems of the form y = Ax⋆ + σyw, where the unknown image x⋆ is a realization of x ∼ p(x), w is additive Gaussian noise, and A and σy are known at inference time. To handle varying scales, BAM conditions on a rescaled measurement,
defined through the stochastic interpolant yσ = αsAx + ςσw, with schedule αs and ςσ satisfying ςσ/αs = σ. This allows the model to generalize across operators and noise levels encountered in practice.
Training Objectives
BAM is trained on two main flow map training objectives:
-
On the diagonal (t = s), the velocity is fitted to the interpolant slope by minimizing Lb(θ) = E[vθ(xt, t, t, yσ, A) − (z − x0)2/2].
-
Off-diagonal (s < t), backwards jumps from t to s are governed by the Lagrangian condition of section 2: LLSD(θ) = E[∂sXθt,s − sg vθ−(xˆt,s, s, s, yσ, A)2/2].
The training objective is L(θ) = E[λbLb + λLLLSD + g(s)λplLPIPS(xˆt,s, x0) + λclctr(xˆt,s, x0)], where g(s) = exp(−4s). The auxiliary losses target perceptual quality and contrast bias:
**)&lLPIPS is the squeeze-based perceptual loss of Zhang et al. (2018), applied to the clipped output. It is weighted by g(s) = exp(−4s), assigning weight only when s is small. **
**)&lctr compares intensity histograms to penalize contrast bias, evaluated before clipping and set to zero after pre-training. **
Key Contributions and Performance
BAM's contributions include:
-
A flow-map training paradigm:
We propose a recipe that upgrades an unfolded reconstruction network into such a flow map, via Lagrangian self-distillation on a normalized stochastic interpolant of the measurement.
-
A lightweight foundation model: Built on the 36M-parameter RAM backbone, it learns this map across operators, noise levels and datasets,
delivering high-quality posterior samples (Figure 1) and generalizing to new data and tasks.
-
State-of-the-art quality at a fraction of the cost:
BAM outperforms specialized models and leading zero-shot methods in sample quality in just three steps, at a fraction of their computational cost.
Experimental Results
Across linear inverse problems on FFHQ, AFHQ, LSUN, DIV2K and the Köhler camera-shake benchmark, BAM outperforms specialized models and leading zero-shot methods in sample quality in just 3 steps both specialised models and leading zero-shot methods in sample quality. Furthermore, BAM⋆ advances the Pareto frontier of the baselines, attaining better quality at a lower cost.
BAM is also robust to post-training quantization under INT8, achieving a 3.51× end-to-end speedup
with only a moderate loss in accuracy. BAM also provides a spatial uncertainty map at no extra training cost for sparse-view CT problems.
Limitations
BAM has four main limitations:
-
It handles only additive Gaussian noise and linear or mildly nonlinear forward operators.
-
It is non-blind, meaning the instrument model must be supplied, so blind problems rely on an external estimate of the operator.
-
It does not detect model misspecification, so under strong distribution shift it can return unreliable posteriors without warning.
-
Posterior sample quality was assessed empirically, and formal guarantees for the learned conditional flow map remain open.
Improvements for AI systems
As a fastidious and diligent researcher, I have analyzed the provided paper, BAM! BAYESIAN ANYTHING MODEL: A FOUNDATION MODEL FOR GENERATIVE COMPUTATIONAL IMAGING.
The core contribution is a lightweight foundation model that enables few-step, physics-aware posterior sampling for general inverse problems.
Here are the specific improvements to AI systems and what those improved systems can achieve:
)
)
-
[Mechanism: Few-Step Posterior Sampling via Conditional Flow Map]
-
[Application: High-Fidelity, Physics-Aware Image Restoration and Reconstruction]
-
[Capability 1: Zero-Shot/Minimal Fine-Tuning Generalization across Operators and Noise Levels]
-
[Capability 2: Efficient Uncertainty Quantification (Spatial Variation Mapping)]
-
[Capability 3: Robustness to Real-World Acquisition Physics (Motion Blur, CT)]
Detailed breakdown of improvements:
- [Mechanism: Few-Step Posterior Sampling via Conditional Flow Map]
A system built on BAM can perform complex image reconstruction by using only a few network evaluations (e.g., 3 steps) instead of hundreds required by traditional diffusion models or consistency models. This is achieved by training BAM as a conditional flow map that transports Gaussian reference distributions directly to the desired posterior distribution, bypassing the need for expensive likelihood approximations or guidance weights during inference.
- [Application: High-Fidelity, Physics-Aware Image Restoration and Reconstruction]
The system can accurately restore images from various degradation types (Gaussian blur, JPEG compression/denoising) and perform complex inverse problems like blind motion deblurring and sparse-view CT reconstruction. Unlike current methods that often hallucinate content or produce overly smooth results, BAM generates posterior samples that preserve sharp textures (e.g., fur, hair) while remaining consistent with the known measurement physics.
- [Capability 1: Zero-Shot/Minimal Fine-Tuning Generalization across Operators and Noise Levels]
The improved system is a foundation model
that generalizes robustly to unseen data and tasks (unseen operators, unseen noise levels). Because it is pre-trained jointly on large image corpora and libraries of forward operators, it can be deployed for new imaging problems immediately with only zero-shot inference or minimal fine-tuning. This drastically lowers the barrier to entry for applying generative models to specialized domains.
- [Capability 2: Efficient Uncertainty Quantification (Spatial Variation Mapping)]
The system provides a novel method for generating a spatial uncertainty map (using the posterior standard deviation calculated across multiple draws) directly from the model output, without requiring extra training costs. For problems like CT reconstruction, this map can highlight anatomical edges and vessels with high precision, offering a quantifiable measure of reliability that traditional point estimators lack.
- [Capability 3: Robustness to Real-World Acquisition Physics (Motion Blur, CT)]
The system is specifically designed to handle complex physical models at inference time. It can utilize an estimated forward operator derived from external tools (like Kernel Prediction Networks for motion blur) or even be adapted for single-channel modalities like CT scans where the acquisition physics is known but the data distribution differs significantly from natural images. This allows it to solve blind
problems where the true physical operator is unknown during training.
This BAM-based system can serve as a highly efficient, versatile generative engine for any task involving image estimation under known (or estimated) physical constraints, moving Bayesian computational imaging from a specialized, dataset-bound endeavor to a general-purpose tool for scientific and medical imaging.
Sources
- Stochastic Interpolants: A Unifying Framework for Flows and Diffusions
- How to build a consistency model: Learning flow maps via self-distillation
- Blind Motion Deblurring with Pixel-Wise Kernel Estimation via Kernel Prediction Networks
- StarGAN v2: Diverse Image Synthesis for Multiple Domains
- Diffusion Posterior Sampling for General Noisy Inverse Problems
- InvFusion: Bridging Supervised and Zero-shot Diffusion for Inverse Problems
- Zero-Shot Image Restoration Using Few-Step Guidance of Consistency Models (and Beyond)
- Deep Equilibrium Architectures for Inverse Problems in Imaging
- GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium
- Denoising Diffusion Probabilistic Models
- Rethinking FID: Towards a Better Evaluation Metric for Image Generation
- Plug-and-Play Methods for Integrating Physical and Learned Models in Computational Imaging
- A Style-Based Generator Architecture for Generative Adversarial Networks
- Denoising Diffusion Restoration Models
- Consistency Trajectory Models: Learning Probability Flow ODE Trajectory of Diffusion
- Regularization by Texts for Latent Diffusion Inverse Solvers
- Bayesian imaging using Plug & Play priors: when Langevin meets Tweedie
- Flow Matching for Generative Modeling
- I$^2$SB: Image-to-Image Schr\"odinger Bridge
- Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models