Unifying Generative Models with Path Integrals
Ramon Winterhalder
Università degli Studi di Milano · INFN
cs.LG, hep-ph, stat.ML
Submitted: 2026-08-12
Updated: 2026-08-14
Comments: 51 pages, 4 figures, 4 tables
Code: https://github.com/ramonpeter/generative-path-integrals
Project page: https://iml-wg.github.io/HEPML-LivingReview
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 95/100
The gist: This paper formulates generative modeling as a path integral, unifying flow-based, diffusion-based, variational, and adversarial models as different evaluation principles for a single master action.
Terminology
Summary
This paper formulates generative modeling as a path integral, unifying flow-based, diffusion-based, variational, and adversarial models as different evaluation principles for a single master action. The authors derive this from latent-variable models extended to chains of latent variables, taking the continuum limit to obtain an Onsager–Machlup action. The central result is the master path-integral form for the data density:
pdata (z0) = ∫ dz1 pprior (z1) ∫ Dz exp [−∫01 dt ∥ż(t) − frev (z(t), t)∥2 / (2g2(t))]
where frev is the reverse drift of the generative process and g(t) is the diffusion coefficient.
The paper demonstrates that:
-
Normalizing flows and continuous normalizing flows arise as the deterministic saddle-point limit (g → 0), recovering the change-of-variables formula and the continuous normalizing-flow log-likelihood.
-
Diffusion models correspond to the stochastic regime, with score matching derived from the KL divergence between transition kernels.
-
Conditional flow matching replaces stochastic sampling by a probability-flow ODE.
-
Schrödinger bridges appear as a global variational principle on the path measure.
-
Variational autoencoders are the degenerate single-latent-variable case with a variational approximation.
-
Generative adversarial networks arise when the likelihood is unavailable (e.g., singular decoder distributions), leading to likelihood-free saddle points.
The paper then recasts the action in Martin-Siggia-Rose-Janssen-de Dominicis (MSRJD) form, separating free (affine) from interacting (nonlinear) probability flows. This enables a diagrammatic perturbation theory where nonlinearities in the drift act as interaction vertices. The expansion yields a one-loop correction to deterministic samplers, consisting of a covariance broadening and a tadpole mean shift, obtained by integrating two auxiliary equations (a Lyapunov equation and a mean-shift equation) alongside the deterministic trajectory. This correction is validated numerically:
-
On a two-dimensional Ornstein–Uhlenbeck free theory, where it is exact, matching stochastic simulations to within 0.14% relative error.
-
On a one-dimensional cubic drift, where it reduces a 53% tree-level error to 1.6% at the largest diffusion strength tested, with residuals scaling as O(g4).
-
On a 24-dimensional permutation- and rotation-equivariant drift, where it reduces errors from 41–83% to 1.1–13% depending on diffusion strength.
The paper also treats imperfect learned scores as diagrammatic insertions, deriving a response-weighted score-matching objective with weighting λresp(t) ∝ g4(t)‖G(0,t)‖2, where G is the response propagator. Finally, it applies effective field theory power counting to construct symmetry-equivariant drifts, enumerating operators up to degree three for SN × O(m) symmetry and predicting a coupling hierarchy controlled by the ratio of latent size to curvature scale.
Improvements for AI systems
Improvements to AI systems based on this paper:
-
Unified generative model selector: Build an AI system that automatically chooses the optimal generative paradigm (flow, diffusion, VAE, GAN, or Schrödinger bridge) for a given dataset by evaluating the master path-integral action. This system can switch between deterministic and stochastic regimes based on the measured diffusion coefficient, avoiding the current trial-and-error selection of model families.
-
One-loop corrected samplers: Implement a new sampling algorithm that augments any deterministic flow or ODE-based generator (e.g., continuous normalizing flows, probability-flow ODEs) with the one-loop correction—integrating the Lyapunov equation for covariance broadening and the mean-shift equation alongside the trajectory. This yields a sampler that is dramatically more accurate for strongly nonlinear drifts (reducing errors from 53% to 1.6% in cubic-drift tests) without the computational cost of full stochastic simulation.
-
Response-weighted score matching: Replace standard score-matching objectives with the derived response-weighted loss, using weight λresp(t) ∝ g4(t)‖G(0,t)‖2. This produces diffusion models that are robust to imperfectly learned scores, improving training stability and final sample quality, especially in low-data or high-noise regimes.
-
Symmetry-aware architecture generator: Use the effective field theory power counting to automatically construct neural network drift architectures with guaranteed SN × O(m) equivariance, enumerating operators up to degree three. This eliminates manual feature engineering for permutation- and rotation-invariant data (e.g., particle physics, point clouds, molecular conformations), and predicts a coupling hierarchy that guides layer initialization for faster convergence.
-
Diagrammatic uncertainty estimator: Add a post-hoc correction module to any generative model that computes the tadpole mean shift and covariance broadening from the learned drift, providing a quantitative measure of how far the model is from the true path measure. This enables the system to flag unreliable samples or regions of latent space where the deterministic approximation fails, improving confidence calibration in downstream tasks.
-
Path-integral-based likelihood estimation: For models where the likelihood is intractable (e.g., GANs with singular decoders), use the MSRJD action to compute an approximate likelihood via the saddle-point plus one-loop expansion. This allows the AI system to perform likelihood-based tasks (anomaly detection, model comparison, Bayesian inference) on top of likelihood-free trained generators, bridging the gap between GANs and probabilistic models.
Abstract
We formulate generative modeling as a path integral in which flow-based, diffusion-based, variational, and adversarial models arise as different evaluation principles for a single master action. Its Martin-Siggia-Rose-Janssen-de Dominicis (MSRJD) form separates free from interacting probability flows and opens them to diagrammatic perturbation theory. The expansion yields a one-loop correction to deterministic samplers at no stochastic-sampling cost, which we validate on solvable and nonlinear drifts, where it reduces a 53 % tree-level error to 1.6 %. Imperfect learned scores enter as insertions and yield a response-weighted score-matching objective, and symmetry-equivariant drift design becomes an operator expansion with EFT power counting.
Sources
- Machine Learning and LHC Event Generation
- Modern Machine Learning for LHC Physicists
- CaloChallenge 2022: A Community Challenge for Fast Calorimeter Simulation
- A Living Review of Machine Learning for Particle Physics
- The Living Guide of Machine Learning for Particle Physics
- Score-Based Generative Modeling through Stochastic Differential Equations
- Flow Matching for Generative Modeling
- Building Normalizing Flows with Stochastic Interpolants
- Diffusion Schr\"odinger Bridge with Applications to Score-Based Generative Modeling
- An optimal control perspective on diffusion-based generative modeling
- SurVAE Flows: Surjections to Bridge the Gap between VAEs and Flows
- Event Generation and Density Estimation with Surjective Normalizing Flows
- A Variational Perspective on Diffusion-Based Generative Models and Score Matching
- Generative Diffusion From An Action Principle
- Understanding Diffusion Models by Feynman's Path Integral
- Auto-Encoding Variational Bayes
- Symmetries of generating functionals of Langevin processes with colored multiplicative noise
- A Guide to Constraining Effective Field Theories with Machine Learning
- Maximum Likelihood Training of Score-Based Diffusion Models
- Understanding Diffusion Objectives as the ELBO with Simple Data Augmentation
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks