Optimizing Likelihoods via Mutual Information: Bridging Simulation-Based Inference and Bayesian Optimal Experimental Design
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Optimizing Likelihoods via Mutual Information: Bridging Simulation-Based Inference and Bayesian Optimal Experimental Design".
Jane: The paper was written by Vincent D. Zaballa and Elliot E. Hui from University of California, Irvine.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back to the show, everyone. Today we're digging into a paper with a real mouthful of a title: "Optimizing Likelihoods via Mutual Information: Bridging Simulation-Based Inference and Bayesian Optimal Experimental Design." Jane, I'm going to need you to break that down for me before my brain melts.
Jane: Happy to, Tom. So there are two big ideas in that title. First, simulation-based inference — that's when you have a computer model of some real-world process, like how a disease spreads or how cells signal each other, and you want to figure out the hidden parameters just by running simulations. The second is Bayesian optimal experimental design — that's figuring out which experiment to actually run to learn the most from your limited resources.
Tom: Right, so one is about learning from data you already have, and the other is about choosing what data to collect. And this paper says those two things are actually the same problem?
Jane: Exactly. They show that the math you use to pick a good experiment is the same math you use to train your inference model. It's like realizing you've been carrying two separate toolboxes when one set of tools does both jobs.
Tom: And that matters because simulators are expensive. Every time you run one, it might take seconds or minutes. So if you can make each run count for both training and design, that's huge.
Jane: That's the core insight. The authors are from UC Irvine, and they've built a method called SBI-BOED that does both simultaneously. They're not just theorizing either — they test it on real biological models, like the BMP signaling pathway that's important in development and disease.
Tom: The BMP model sounds like something we should come back to. But first, tell me why this connection between inference and design is so surprising.
Jane: Well, traditionally these fields developed separately. Inference folks cared about getting accurate posteriors from fixed data. Design folks cared about maximizing information gain. This paper shows that if you use the right objective — a mutual information bound — you're actually training your likelihood model and optimizing your experiment at the same time.
Tom: So it's not just a clever trick, it's a fundamental link between two fields that thought they were doing different things.
Jane: Exactly. And that's what makes this paper exciting. It's not just a new algorithm, it's a new way of thinking about the problem.
Tom: I'm sold already. Let's get into the actual method and how they made this work in practice.
Summary: Tom: So we've established the big idea — connecting inference and design through mutual information. But how did they actually pull this off? What's the concrete method?
Jane: They use something called InfoNCE, which is a way to estimate mutual information using contrastive samples. You draw a bunch of parameter values, simulate data for one of them, and then ask your model to pick out which parameter generated that data. The better it does, the more information you're capturing.
Tom: And the trick here is that they use a normalizing flow as their likelihood model. That's a type of neural network that gives you a proper probability distribution. So instead of just a classifier that says "this looks right," you get an actual likelihood function you can use for downstream inference.
Jane: Right. And here's where it gets interesting. They add a regularization term — they call it lambda — that controls the tradeoff between maximizing information gain and fitting the likelihood accurately. It's like a dial between "explore new experiments" and "learn the model well."
Tom: And they found that dial matters a lot, right? In their Two Moons benchmark, they showed that cranking up the regularization improves calibration but lowers the information gain estimate.
Jane: Exactly. It's a genuine tradeoff. But the key result is that their method, SBI-BOED, works even when the simulator is a closed box — you can't differentiate through it. That's huge because most real scientific simulators are like that. You can run them, but you can't get gradients from them.
Tom: That's the part that got me excited. Previous methods like iDAD required differentiable simulators. This one doesn't. So it opens up a whole class of problems that were previously out of reach.
Jane: And they show it works on two real models. The SIR epidemiology model — susceptible, infected, recovered — and the BMP signaling pathway. For SIR, they got better posterior accuracy than iDAD while using about fifty times fewer simulator calls.
Tom: Fifty times fewer. That's not incremental, that's transformative. For anyone running expensive simulations, that's the difference between a project being feasible or not.
Jane: And on BMP, which is a closed-box simulator, they beat the strong baseline on calibration and predictive accuracy even though the baseline had a higher estimated information gain.
Tom: Wait, that's counterintuitive. How can you have lower information gain but better predictions?
Jane: That's one of the most interesting findings. The information gain metric can be misleading. A design that scores high on paper might not give you a posterior that's actually well-calibrated. So they're arguing we need to look beyond just the EIG number.
Tom: That's a provocative claim. It challenges how the whole field evaluates experimental design methods.
Jane: It does. And it's backed by real experiments. That's what makes this paper worth taking seriously.
Improvements: Tom: We're back with the paper "Optimizing Likelihoods via Mutual Information: Bridging Simulation-Based Inference and Bayesian Optimal Experimental Design." Jane, you mentioned they found EIG can be misleading. What improvements do they actually propose to fix these issues?
Jane: Well, they tackle several practical problems. First, there's the issue of sparse rewards — when you're optimizing designs, sometimes the gradient signal is just flat, so your optimizer gets stuck. They solve that by optimizing a distribution over designs instead of a single design. It's like exploring a whole neighborhood of experiments rather than just one point.
Tom: So instead of asking "what's the best design," you ask "what's the best region of designs to sample from." That gives you more stability.
Jane: Exactly. They use a truncated normal distribution that starts wide and narrows over time. That way, even if the reward landscape is bumpy, the distribution has enough support to find good regions.
Tom: And they also use checkpoints, right? Like in deep learning where you save the best model during training.
Jane: Yes. They checkpoint the design that achieved the highest EIG during training, so even if the optimizer wanders into a bad region at the end, they keep the best design they found. It's a simple trick but it prevents catastrophic forgetting of good designs.
Tom: And then there's the active learning component. They don't just optimize designs — they also decide which simulator calls to make when designs are fixed.
Jane: That's the EPIG part — expected predictive information gain. They use it to prioritize which parameter values to simulate next. Instead of drawing parameters uniformly from the prior, they pick ones where the model is most uncertain. They approximate that uncertainty using MC-dropout, which is a cheap way to get epistemic uncertainty from a neural network.
Tom: So they're being smart about both the design and the training data. Every simulator call is chosen to be maximally informative.
Jane: Right. And the results show this active learning approach improves calibration faster than random sampling under the same simulation budget. It's a complete package — design optimization, likelihood training, and simulation allocation all driven by the same information-theoretic principle.
Tom: That's elegant. One principle driving everything. But I want to push on something — they also mention sequential rounds of inference. How does that fit in?
Jane: They show that you can refine your likelihood between design rounds using the observed data, like traditional SBI methods do. This improves calibration over multiple rounds. It's a modular approach — you can add refinement on top of the design optimization.
Tom: So it's not just a single-shot method, it's a framework that can be extended and improved.
Jane: Exactly. And that flexibility is what makes it practical for real scientific workflows where you might have multiple rounds of experiments.
Conclusion: Tom: Alright, we're wrapping up our discussion of "Optimizing Likelihoods via Mutual Information: Bridging Simulation-Based Inference and Bayesian Optimal Experimental Design." Jane, give us the final takeaway.
Jane: The big picture is that this paper unifies two fields that were working in parallel. By showing that mutual information maximization is equivalent to likelihood learning, they've given us a single objective that handles both inference and experimental design. And they've made it work for closed-box simulators, which is where most real science happens.
Tom: And the practical impact is real. Fifty times fewer simulator calls on SIR, better calibration on BMP, and a method that doesn't require differentiability. That's going to change how people run experiments in biology, epidemiology, any field with expensive simulations.
Jane: But I think the most important contribution is the cautionary note — that EIG alone isn't enough. They showed that higher information gain doesn't always mean better posteriors. That's a wake-up call for the whole BOED community.
Tom: It's a humbling reminder that our metrics can lie to us. But it's also exciting because it opens up new research directions — how do we design experiments that are both informative and produce calibrated models?
Jane: And the framework they've built is flexible enough to incorporate new generative models, like diffusion models, as they become more practical. This isn't the end of the story, it's the beginning.
Tom: Well said. Thanks to everyone who joined us today — Lu, Meng, Lalam, and all our listeners. We'll be back with another paper soon. Until then, keep questioning your metrics and stay curious.
Jane: See you next time, everyone.
Vincent D. Zaballa, Elliot E. Hui
University of California, Irvine
stat.ML, cs.LG
Submitted: 2026-08-13
Updated: 2026-08-14
Comments: Accepted at the 42nd Conference on Uncertainty in Artificial Intelligence (UAI 2026)
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 63/100
Key concepts
- Simulation-Based Inference (SBI)
- A method used to determine hidden parameters of a real-world process by running computer simulations. Instead of traditional data, the inference model learns directly from the outputs generated by running the simulation multiple times.
- Bayesian Optimal Experimental Design (BOED)
- The process of deciding which experiment to run next to maximize learning from limited resources. The goal is to choose a design that yields the most information, thereby improving parameter estimates efficiently.
- Mutual Information
- A mathematical measure used in the paper's objective function. Maximizing mutual information serves as a single principle that simultaneously trains the likelihood model and optimizes the experimental design.
- Closed-Box Simulator
- A type of real scientific simulator that can be run to generate data but does not allow for mathematical differentiation (gradients). The paper's method is significant because it works with these simulators, which are common in real science.
Terminology
Summary
Summary
This paper introduces SBI-BOED, a unified framework that bridges Simulation-Based Inference (SBI) and Bayesian Optimal Experimental Design (BOED) through a shared mutual information (MI) perspective. The authors demonstrate that the InfoNCE lower bound on Expected Information Gain (EIG) serves as a principled training objective for conditional normalizing flow density models—the standard workhorse of SBI—with an asymptotic likelihood-fitting interpretation. This connection allows a single stochastic-gradient procedure to jointly train a surrogate likelihood and optimize experimental designs without requiring simulator differentiability or pathwise gradients to designs.
Core Theoretical Contribution: The paper establishes Proposition 3.1, showing that "in the limit as the number of contrastive samples L → ∞, maximizing the lower bound of the Mutual Information (MI) between parameters θ and observations y in a Simulation-Based Inference (SBI) setting is equivalent, up to constants independent of ϕ, to minimizing the Kullback-Leibler (KL) divergence, D KL(p(yθ)p ϕ(yθ)), between the likelihood p(yθ) and its approximation p ϕ(yθ), together with the corresponding marginal-likelihood approximation term." The proof shows that maximizing the InfoNCE bound asymptotically equals optimizing the likelihood objective with a finite-L marginal-likelihood approximation term whose accuracy depends on the number of contrastive samples.
Regularized MI Bound (InfoNCE-λ): The authors propose a regularized objective:
L NCE-λ(ξ, ϕ; L):= E[log(p ϕ(yθ0, ξ) / ((1/(1+L)) Σ l=0 L p ϕ(yθ l, ξ))) + λ · log p ϕ(yθ0, ξ)],
where the expectation is over p(θ0)p(yθ0,ξ)p(θ1:L). Theoretical analysis shows that positive λ reduces the weight of the approximate likelihood term (1−λ) to ensure broad coverage of parameter space, while negative λ amplifies likelihood weight leading to mode-seeking behavior. The EIG estimate linearly depends on λ: Positive λ decrease the EIG estimate while negative values increase the EIG at the expense of likelihood approximation.
Active Learning via EPIG: For fixed designs, the same MI-based objective induces an active learning strategy over simulator parameters. The authors use MC-dropout to form stochastic likelihood particles
ϕ m and estimate Expected Predictive Information Gain (EPIG) as:
EPIG(θ) ≜ E[log(p ϕ(y, y⋆ θ, θ⋆, ξ) / (p ϕ(y θ, ξ) p ϕ(y⋆ θ⋆, ξ)))],
selecting θ with maximal EPIG to prioritize simulator queries. This makes simulator calls more effective even in the absence of design optimization.
Design Distributions and Checkpoints: To address sparse-reward landscapes, the authors optimize a truncated Normal distribution over designs p ψ(ξ) = N trunc(μ ξ σ n2) rather than point designs, with an exponential decay schedule σ n = σ end + (σ start − σ end)·e(−n*ρ/N). They also employ design checkpoints to save designs achieving high EIG during training, mitigating local optima issues.
Experimental Results:
-
Two Moons (fixed-design SBI): Sweeping λ and L shows
increasing L improves the MI lower bound and generally improves likelihood validation metrics, while λ controls the tradeoff between emphasizing the MI objective and likelihood accuracy.
EPIG-based active learning yieldsmore efficient improvement in calibration over rounds under the same simulation budget
compared to random θ sampling. Larger λ worsens local calibration (higher l-C2ST), while K (top-K contrastive negatives) acts as a variance-control knob with non-monotone effects. -
Noisy Linear Model: SBI-BOED optimizes EIG across design dimensions (1D, 10D, 100D), with λ regularization stabilizing training in high-dimensional design spaces at
the cost of slightly lower EIG.
The 100-dimensional posterior concentrates near true parameters relative to the broad prior. -
SIR Model (T=2 rounds): SBI-BOED achieves EIG of 1.01–1.63 (depending on λ) versus 2.69 for MINEBED and 2.67 for iDAD, but significantly outperforms baselines on posterior quality: L-C2ST of 0.03–0.05 versus 0.07–0.10 for baselines, and median distance of 46.85–52.85 versus 71.26–71.65 for baselines. The authors note
improved information gain does not necessarily correlate with improved downstream prediction.
The ablation study shows that without a design distribution, optimizationfails to find designs with high rewards
in the first round. -
BMP Model (T=3 rounds): On this non-differentiable simulator, SBI-BOED achieves EIG of 10.38–10.39 versus 14.23 for MINEBED-BO, but with dramatically better posterior metrics: L-C2ST of 0.01 versus 0.25, and median distance of 0.60–0.61 versus 0.82. Wall-clock time was 2,235 ± 9 s for SBI-BOED versus 7,095 ± 2,519 s for MINEBED-BO.
Simulation Efficiency: Table 1 shows SBI-BOED uses approximately 5.1 × 106 simulator calls on the SIR model, versus approximately 2.6 × 108 for iDAD and 107–108 for RL-BOED—a 50× reduction in simulator calls while exceeding iDAD accuracy.
Sequential SBI Refinement: The authors demonstrate that incorporating SBI-style iterative inference (Algorithm 2) after each BOED round improves posterior calibration. For the BMP model, refinement over multiple rounds yields calibrated posteriors, showing the modularity and flexibility in improving posterior inference that SBI methods offer in between BOED rounds.
Key Findings and Caveats: The paper emphasizes that designs with higher EIG do not necessarily yield better downstream inference, such as posterior calibration and predictive accuracy diverging from information-gain objectives.
The authors also note that MI optimization does not eliminate mode collapse in Two Moons likelihood-based models, attributing this to deficiencies in the marginal likelihood approximation
from finite contrastive samples and approximate likelihoods. The approach extends to any likelihood-based generative model supporting log p ϕ(yθ,ξ) evaluation, with natural extensions to diffusion models and flow-matching suggested.
Improvements for AI systems
Based on the paper, here are the specific improvements I can implement in AI systems:
Improvement: Replace standard maximum-likelihood training of conditional normalizing flows with the InfoNCE-λ objective (Equation 9) that jointly optimizes mutual information and likelihood fidelity.
What the improved system can do:
-
Train likelihood surrogates on 50× fewer simulator calls while maintaining or improving posterior accuracy (demonstrated on SIR model: 5.1×106 vs 2.6×108 calls for iDAD)
-
Automatically balance information-seeking behavior against likelihood accuracy via the λ hyperparameter
-
Avoid mode collapse in multimodal posteriors (e.g., Two Moons benchmark) by controlling the tradeoff between MI maximization and likelihood fitting
Abstract
Simulation-based inference (SBI) is a method to perform inference on a variety of complex scientific models with challenging inference (inverse) problems. Bayesian Optimal Experimental Design (BOED) aims to efficiently use experimental resources to make better inferences. Various stochastic gradient-based BOED methods have been proposed as an alternative to Bayesian optimization and other experimental design heuristics to maximize information gain from an experiment. We demonstrate a link via mutual information bounds between SBI and stochastic gradient-based variational inference methods that permits BOED to be used in SBI applications as SBI-BOED. This link allows simultaneous optimization of experimental designs and optimization of amortized inference functions. We evaluate the pitfalls of naive design optimization using this method in a standard SBI task and demonstrate the utility of a well-chosen design distribution in BOED. We compare this approach on SBI-based models in real-world simulators in epidemiology and biology, showing notable improvements in inference.
Sources
- Optimizing Sequential Experimental Design with Deep Reinforcement Learning
- Statistically Efficient Bayesian Sequential Experiment Design via Reinforcement Learning with Cross-Entropy Estimators
- Maximum Likelihood Learning of Unnormalized Models for Simulation-Based Inference
- Active Sequential Posterior Estimation for Sample-Efficient Simulation-Based Inference
- Bayesian Experimental Design for Implicit Models by Mutual Information Neural Estimation
- Gradient-based Bayesian Experimental Design for Implicit Models using Mutual Information Lower Bounds
- Probabilistic Bayesian optimal experimental design using conditional normalizing flows
- Prediction-Oriented Bayesian Active Learning
Related papers
- Behavior of prediction performance metrics with rare events
- Optimal Estimation of Generic Dynamics by Path-Dependent Neural Jump ODEs
- A Posterior-Dynamics Framework for Imaging Inverse Problems with Pretrained Diffusion Priors
- One Permutation Is All You Need: Fast, Deterministic Feature Importance and Model Stress-Testing
- Online Conformal Prediction for Non-Exchangeable Panel Data
- Deep Time-Series Forecasting in 10 Years: A Survey