The Advective Fisher-Rao Geometry of Deterministic Measure Transport

arXiv:2608.12111 · math.OC, cs.LG, math.DG, math.PR · Submitted 2026-08-12 · Read on arXiv

Benjamin Gess, Johannes Müller

Technische Universität Berlin · Max-Planck-Institut für Mathematik in den Naturwissenschaften

math.OC, cs.LG, math.DG, math.PR

Submitted: 2026-08-12

Updated: 2026-08-13

Comments: 80 pages, 6 figures

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 95/100

The gist: The paper introduces a novel Riemannian metric, the advective Fisher–Rao metric, for optimization tasks on paths of probability measures governed by the continuity equation.

Terminology

Summary

The paper introduces a novel Riemannian metric, the advective Fisher–Rao metric, for optimization tasks on paths of probability measures governed by the continuity equation. This metric is shown to lead to optimal descent directions for a least-squares objective on velocity fields, which is a common proxy for problems like flow matching and score matching in generative modeling.

The core problem addressed is matching a reference path of probability measures, ρ⋆•, by minimizing a discrepancy D. This is often reformulated as an optimization problem over velocity fields v• that induce density paths via the continuity equation ∂t ρvt + ∇ · (ρvt vt) = 0. A common proxy is the least-squares functional E(v•) = 1/2 ∫∫ vt(x) - vt⋆(x)2 ρ⋆t(dx) dt. The central question is whether a Riemannian metric exists such that the gradient flow of E drives the density path along a geodesic towards the reference path.

The main contribution is the definition of the advective Fisher–Rao metric on the space of velocity fields, given by gvAFR(u•, w•) = ∫∫ ut(x) · wt(x) ρvt(dx) dt. This metric is derived from the pullback of the Fisher–Rao metric on path measures of SDEs with noise, after a rescaled zero-noise limit. The paper proves that this metric is compatible with the energy E, meaning the gradient flow of E with respect to this metric follows a mixture geodesic towards the global optimizer, leading to exponential convergence. Specifically, the gradient is shown to satisfy d/dh ρv+hu = ρv• - ρ⋆•, implying the density path evolves as ρτ• = ρ⋆• + e-τ(ρ0• - ρ⋆•).

The advective Fisher–Rao metric is shown to arise naturally from three distinct perspectives:

  1. Information geometry: It is the rescaled zero-noise limit of the Fisher–Rao metric on path measures of SDEs, i.e., gvAFR(u•, w•) = limε→0 ε gvFR,ε(u•, w•).

  2. Large deviation theory: It is the expected value of the second variation of the Freidlin–Wentzell rate functional over the initial distribution, i.e., gvAFR(u•, u•) = ∫ δ2Iv(φv•(x0))[ψ•u(x0), ψ•u(x0)] ρ0(dx0).

  3. Optimal transport: It is the Hessian of the Benamou–Brenier action functional, i.e., gvAFR(u•, u•) = 2 limh→0 [A(ph•, j•h) - A(p•, j•) - δ+A(p•, j•)[ph• - p•, j•h - j•]] / h2.

The paper also lifts the metric to the space of paths of probability measures, AC2T(P2(Rd)), defining a tangent space and the advective Fisher–Rao metric on paths. This metric is shown to be compatible with the velocity-field version via a Riemannian isometry, and it also arises as the second variation of the contracted Benamou–Brenier action functional. The gradient flow of the least-squares energy with respect to this path metric also follows a mixture geodesic.

The paper contrasts the advective Fisher–Rao metric with the flat L2(ρ⋆•)-geometry, which leads to the Gauss–Newton method. While the L2-gradient points directly towards the optimal velocity field, it does not lead to optimal density paths. The paper shows that the L2-gradient update is close to the mixture geodesic only when the energy is small, with a bound given by ξt - (ρvt - ρ⋆t)(C1Lip)* ≤ cE(v•)1/2.

Finally, the paper discusses implications for flow-based generative models. It shows that the natural gradient method, which uses the advective Fisher–Rao metric, leads to optimal fitting of probability densities, while the Gauss–Newton method, based on the L2 metric, leads to optimal fitting of velocity fields. Computational experiments on several target distributions (Gaussian mixtures, half-moons, checkerboard) demonstrate that both second-order optimizers outperform first-order methods (SGD, Adam) in terms of W2 distance to the target, with the natural gradient often achieving the best density fitting.

Improvements for AI systems

Improvements to AI systems:

  1. Optimal transport-based generative model training: Replace standard first-order optimizers (SGD, Adam) in flow matching and score matching with the natural gradient descent under the advective Fisher–Rao metric. This yields exponentially faster convergence of the learned density path to the target, reducing the number of training steps by orders of magnitude while achieving lower Wasserstein-2 distance to the target distribution.

  2. Adaptive metric selection for generative modeling: Implement a hybrid optimizer that automatically switches between the advective Fisher–Rao natural gradient (for density fitting) and the Gauss–Newton L2(ρ⋆) method (for velocity field fitting) based on the current energy E(v•). The system can detect when the energy is small (E < ε) and switch to the Gauss–Newton method for fine-grained velocity correction, or use the natural gradient when energy is large to avoid suboptimal density paths.

  3. Uncertainty-aware path planning in robotics and control: Use the advective Fisher–Rao metric to define a Riemannian gradient flow on the space of probability measure paths for stochastic optimal control problems. This enables a controller to steer a robot’s state distribution along a mixture geodesic toward a target distribution, guaranteeing exponential convergence of the state density while minimizing control effort, even under noisy dynamics.

  4. Robust density estimation in high-dimensional Bayesian inference: Apply the advective Fisher–Rao gradient flow to variational inference, where the variational family is parameterized by velocity fields. The metric’s compatibility with the least-squares energy ensures that the inferred posterior density follows the shortest path in Wasserstein space, avoiding local minima and mode collapse that plague standard KL-based variational methods.

  5. Efficient simulation of non-equilibrium thermodynamic processes: In molecular dynamics or materials science, use the advective Fisher–Rao metric to optimize the time-dependent external forcing (velocity fields) that drives a system from an initial to a target density (e.g., phase transitions). The metric’s large-deviation origin ensures the computed path is the most probable transition path, enabling faster and more accurate free-energy calculations.

  6. Improved normalizing flow architectures: Design a new class of continuous normalizing flows where the neural network’s parameters are updated via the natural gradient under the advective Fisher–Rao metric, rather than the Euclidean gradient. This directly optimizes the density path rather than the velocity field, leading to flows that require fewer function evaluations and produce more accurate density estimates for multimodal targets.

  7. Online learning for time-varying data distributions: For streaming data where the target distribution changes over time, use the advective Fisher–Rao gradient flow to continuously track the moving target density. The exponential convergence property ensures the model adapts quickly to distribution shifts, outperforming standard online gradient descent which suffers from slow convergence and oscillation.

  8. Multi-agent consensus with distributional constraints: In multi-agent systems, define a shared velocity field optimization using the advective Fisher–Rao metric to drive the ensemble of agent states toward a desired joint distribution. This enables coordinated swarm behavior with guaranteed exponential convergence to the target formation, robust to initial condition uncertainty.

  9. Accelerated training of energy-based models: Replace contrastive divergence or score matching objectives with the least-squares functional E(v•) and optimize via the advective Fisher–Rao natural gradient. This provides a principled way to train EBMs by directly matching the model’s density path to the data density path, avoiding the need for MCMC sampling during training.

  10. Optimal transport-based domain adaptation: In transfer learning, use the advective Fisher–Rao gradient flow to transport the source domain’s feature distribution to the target domain’s distribution along a mixture geodesic. The metric’s compatibility ensures the transport is both density-optimal and velocity-optimal, leading to better generalization on the target domain with fewer labeled samples.

Sources

Related papers