Beckmann Transport Models: From Autonomous Flows to One-Step Maps

arXiv:2608.01692 · cs.LG · Submitted 2026-08-13 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Beckmann Transport Models: From Autonomous Flows to One-Step Maps".

Jane: The paper was written by Lee Cheuk-Kit, Florentin Coeurdoux, Yuyuan Chen, Sophia Tang, Peter Potaptchik et al. from Harvard University and Capital Fund Management and University of Oxford and New York University and University of Pennsylvania.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title: Tom: Welcome back to the arXiv channel, everyone. I'm Tom, and with me as always is Jane. We are looking at a paper that has one of those names that sounds intimidating until you break it down — "Beckmann Transport Models: From Autonomous Flows to One-Step Maps."

Jane: And Tom, I have to say, the title actually tells you a lot. Beckmann is a name from transportation theory, and the phrase "one-step maps" is the part that should get everyone excited, because that's about generating images in a single pass.

Tom: Right, so let's unpack that. In generative modeling, you usually have a process that takes random noise and gradually turns it into a picture, like a time-lapse. This paper is asking whether you can skip the time-lapse and just learn a single function that does the whole job at once.

Jane: Exactly. And the "autonomous flows" part is the clever twist. Normally, the process depends on time — you know exactly where you are in the generation. Here, they're saying, what if the velocity field doesn't depend on time at all? Just a fixed rule for how to move.

Tom: Which sounds like it shouldn't work, right? If you just have a fixed rule, how do you know when to stop? How do you make sure you actually land on the target distribution?

Jane: That's the question the paper answers. And the answer involves something called Beckmann's transportation problem, which is this classical idea from the 1950s about moving mass from one place to another efficiently.

Tom: So we've got a 1950s math problem meeting modern deep learning. I love that combination.

Jane: Me too. And the authors — this is a big team from Harvard, NYU, Oxford, and a few other places — they prove that under the right conditions, this time-independent flow actually does transport your noise distribution exactly to your target distribution.

Tom: And that's a real theorem, not just a heuristic. They're saying, under these geometric conditions, the flow provably works.

Jane: Provably works, and then they go further. Once you have that autonomous flow, you can define a one-step map — a single function that takes a noise point and outputs the final sample. That's the dream for fast generation.

Tom: So the title is really a promise. Beckmann transport gives you the theory, autonomous flows give you the mechanism, and one-step maps give you the speed.

Jane: And we're going to spend this whole episode unpacking how those pieces fit together. Stick around, because the implications for image generation are genuinely exciting.

Summary: Tom: So Jane, we've got the title unpacked. Now let's get into what the paper actually does, because the summary is dense. The core idea is that they're minimizing a flow matching loss, but over time-independent fields instead of time-dependent ones.

Jane: Right, and that's a subtle but huge change. In standard flow matching, you learn a velocity that changes with time — at the start of generation, you move fast, at the end, you slow down. Here, the velocity is the same at every moment.

Tom: And the paper proves this works when the target distribution lives on a lower-dimensional manifold. Think of it like this: your target is a thin sheet or a set of points inside a high-dimensional space.

Jane: That's a great way to put it. And the key requirement is that the target has codimension at least two — meaning it's at least two dimensions smaller than the space it lives in.

Tom: Why does that matter? What breaks if the target is too high-dimensional?

Jane: Because of how the flow behaves near the target. If the target is a surface in three dee space, the flow can pass right through it. But if it's a curve or a point, the flow gets trapped and has to stop there.

Tom: So the geometry is doing the work. The flow naturally converges to the target because there's nowhere else to go.

Jane: Exactly. And they prove this rigorously. They show that almost every trajectory starting from the base distribution ends up on the target manifold, and the distribution of those endpoints matches the target distribution exactly.

Tom: That's the transport property. And then they connect it to Beckmann's problem, which is about finding a flow of mass that satisfies a certain balance equation — the divergence of the flow equals the source minus the sink.

Jane: And that's where the elegance comes in. The flow they construct satisfies exactly that balance equation. The base distribution is the source, the target is the sink, and the flow carries mass from one to the other.

Tom: So it's not just a heuristic that happens to work. It's a flow that satisfies a classical conservation law.

Jane: Precisely. And there's a beautiful payoff. Because the flow is autonomous, the endpoint map — the function that takes a starting point and gives you the final sample — satisfies a simple conservation equation along the flow lines.

Tom: Which means you can learn that map directly, without ever simulating the flow.

Jane: That's the one-step map. And they show that a simple stop-gradient objective can learn it directly from samples. That's the practical breakthrough.

Tom: And they validate it on ImageNet at two hundred fifty-six by two hundred fifty-six resolution, which is no small feat.

Jane: Not at all. We'll get into those results in a moment, but the summary is this: they've built a theoretically grounded framework that connects classical transportation theory to modern generative modeling, and it delivers a one-step map.

Improvements: Tom: So Jane, we've covered the theory. Now let's talk about what this paper actually improves over what came before. Because the authors are pretty direct about fixing a problem in an existing method called Equilibrium Matching.

Jane: Right, and this is where it gets practical. Equilibrium Matching, or EqM, was proposed last year as a way to learn time-independent flows. But the paper points out that the loss used there doesn't enforce the right balance between source and sink.

Tom: Meaning the flow might converge to the target, but with the wrong weights. Some parts of the target get too much mass, others get too little.

Jane: Exactly. And they demonstrate this with a simple experiment. They take a target with five points, each with a specific probability weight, and they show that the original EqM loss gets the weights badly wrong — a twenty-fold increase in error.

Tom: And their fix is elegant. Instead of using a separate schedule to scale the regression target, they tie the target to the actual velocity of the interpolant. That way, the divergence condition is satisfied by construction.

Jane: It's a small change in the loss, but it makes the math work. And on ImageNet, that correction improves the FID score from one point nine zero to one point eight seven, which is modest but consistent.

Tom: Modest on ImageNet, but the bias correction is dramatic on the toy example. So the fix matters, even if the benchmark numbers don't move much.

Jane: And then there's the bigger improvement: the one-step map. They train a network to directly output the final sample, no iterative refinement needed.

Tom: And that's where the numbers get interesting. Their one-step model achieves an FID of seventeen point five eight on ImageNet without any guidance. That's not state-of-the-art, but it's competitive with other one-step methods that don't use guidance.

Jane: Right, and the comparison table is honest about that. Methods that use classifier-free guidance do better, but they also use more compute at inference time.

Tom: So the improvement here is about the architecture of the approach, not just the final number. They're showing that a theoretically grounded one-step map is possible.

Jane: And they also show something practical about iteration. If you apply the learned map multiple times, the samples sharpen. Three applications bring the quality close to what you'd get from the full flow.

Tom: So you can trade compute for quality at inference time, without retraining.

Jane: Exactly. And that's a nice property for deployment. You can ship a fast model and let users choose how many steps they want.

Tom: There's also a training-free version — the Coulomb transport — which is basically a closed-form solution using physics. No neural network needed at all.

Jane: That's the Poisson flow connection. They show that a simple Coulomb field, computed directly from the data, can transport noise to the target. It's slower, but it requires zero training.

Tom: So the improvements span the whole pipeline: a corrected loss, a direct map, an iterable map, and a training-free baseline.

Jane: And that's what makes this paper feel like a complete package, not just a single trick.

Conclusion: Tom: We're wrapping up our discussion of "Beckmann Transport Models: From Autonomous Flows to One-Step Maps," and I have to say, Jane, this one left me genuinely excited.

Jane: Me too, Tom. The paper manages to connect a 1950s transportation problem to cutting-edge generative modeling, and it does so with real theorems, not just heuristics.

Tom: Let's recap what we learned. The core idea is that a time-independent flow, learned with the right loss, can exactly transport a base distribution to a target distribution when the target lives on a sufficiently low-dimensional manifold.

Jane: And that flow satisfies a classical conservation law — the divergence equation from Beckmann's problem. The base is the source, the target is the sink.

Tom: From that flow, they derive a one-step map that can be learned directly. No iterative simulation needed at inference time.

Jane: And they fixed a real bias in Equilibrium Matching, showed that the corrected loss performs better, and demonstrated that the one-step map works on ImageNet at two hundred fifty-six resolution.

Tom: The numbers aren't the best in the field yet, but the framework is solid. And the iterated map gives you a way to trade compute for quality.

Jane: I also appreciated the honesty about limitations. The theory requires the target to have codimension at least two, and the one-step results without guidance are still behind the best guided methods.

Tom: But the path forward is clear. They mention optimizing jointly over the weight and the current, and they hint at applications in text generation, where tokens are naturally discrete points.

Jane: That's a fascinating direction. If you can learn a one-step map to discrete outputs, that could change how we think about language model sampling.

Tom: So we're saying goodbye to this paper, but not to the ideas in it. I suspect we'll see follow-ups soon.

Jane: Absolutely. And with that, we'll wrap up. Thanks for listening, and we'll see you on the next paper.

Tom: Take care, everyone.

Lee Cheuk-Kit, Florentin Coeurdoux, Yuyuan Chen, Sophia Tang, Peter Potaptchik, Yilun Du, Michael S. Albergo, Eric Vanden-Eijnden

Harvard University · Capital Fund Management · University of Oxford · New York University · University of Pennsylvania

cs.LG

Submitted: 2026-08-13

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 73/100

Key concepts

One-Step Map
In generative modeling, this is a single function that takes random noise and directly outputs the final sample in one pass. This contrasts with standard methods that require simulating a time-lapse or flow.
Autonomous Flows
A type of generation process where the velocity field (how the data moves) does not depend on time. This fixed rule for movement is key to creating a fast, single-step output.
Beckmann's Transportation Problem
A classical mathematical idea from the 1950s concerning efficiently moving mass from an initial source distribution to a target distribution. The flow must satisfy this conservation law.
Codimension
A geometric requirement for the target distribution. It means the target is at least two dimensions smaller than the space it lives in, which helps ensure the flow naturally converges to and stays on the target manifold.

Terminology

Summary

Affiliations: Harvard University, Capital Fund Management, University of Oxford, New York University, University of Pennsylvania

Publication: arXiv:2608.01692v3 [cs.LG], 13 Aug 2026


The paper proposes "an instantiation of flow matching that relies on a time-independent velocity field—an autonomous flow—to exactly map between two distributions when the target is supported on a sufficiently lower-dimensional data manifold. The authors also show that the associated one-step generative map is constant in the direction of the autonomous velocity field and therefore satisfies a simple conservation equation, which motivates learning the map directly from samples. These autonomous flows and maps give a dynamical meaning to the flux constraint of Beckmann's transportation problem, and the construction extends beyond flow matching to other probability fluxes satisfying the same constraint. The framework includes a closed-form Coulomb transport closely related to Poisson Flow and a flow-matching-consistent version of Equilibrium Matching. The authors validate their approach on ImageNet 256 × 256."

Standard flow matching and diffusion models drive generation by transporting samples from a base distribution µ0 to a target distribution µ1 using a time-dependent velocity field (or drift) learned by quadratic regression. The paper explores an intriguing alternative: constraining the drift to be time-independent, minimizing the same regression loss over fields b: Rd → Rd that do not depend on t, and performing generation using the autonomous equation Ẋt = b(Xt) solved with initial X0 ∼ µ0. This is essentially the construction proposed by Wang and Du [2025] under the name Equilibrium Matching (EqM).

The authors note that the original EqM paper does not establish this transport property, and its proposed loss need not enforce the required source–sink balance. The paper provides a rigorous justification under explicit geometric and regularity conditions. Specifically, when µ1 is supported on a manifold M1 ⊂ Rd of codimension at least two, the time-independent flow matching drift defines a valid transport: trajectories of the autonomous flow converge to M1, and the resulting map pushes µ0 forward to µ1.

The proof uses the finite FM current and the divergence equation satisfied by the drift. The endpoint map T is selected by following each autonomous trajectory to its endpoint and then freezing it there; it satisfies a simple conservation equation along its characteristics. This equation motivates a stop-gradient objective for training T directly from samples.

The paper's main contributions are:

  1. Proof of valid transport for time-independent FM drift: "Under the regularity and codimension conditions of Assumptions 1–2, the time-independent FM drift defines a valid transport: the minimizer of the FM regression loss over time-independent fields b generates a flow Ẋt = b(Xt) that pushes µ0 to µ1 (Proposition 1). This gives a flow-matching-consistent correction of Equilibrium Matching."

  2. Weighted Beckmann current transport: "For a finite positive weight ν, we show that the unique minimizer of a weighted Beckmann current problem defines a valid autonomous transport when it points uniformly toward the target near M1 (Theorem 1). Gradient-constrained FM provides a natural route to this minimizing current; the canonical Coulomb field with ν ≡ 1 is treated separately (Proposition 2)."

  3. Eulerian equation for the flow map: The map x ↦ Xt(x), which gives the position at time t of the autonomous trajectory starting from x, satisfies the Eulerian equation ∂tXt = b · ∇Xt (Proposition 3).

  4. Conservation equation for the one-step map: "The endpoint of each trajectory defines the one-step map T(x), which pushes µ0 forward to µ1 (Proposition 1). This map is constant in the direction of the autonomous field b, so that b · ∇T = 0, with T = id on M1 (Theorem 2)."

  5. Stop-gradient objective for direct map learning: "The conservation equation b · ∇T = 0 motivates a stop-gradient objective for direct map learning. We prove that the population stop-gradient update vanishes at the true map T (Theorem 3) and evaluate the resulting one-step and iterated inference procedures experimentally."

The paper situates itself relative to several research threads:

Flow matching, diffusion, and equilibrium matching: Flow matching and diffusion models construct differential equations with time-dependent drifts that transport between base and target distributions. The BTM drift minimizes the same FM loss, but over time-independent fields. BTM establishes an exact transport result for a flow-matching-consistent version of Equilibrium Matching and shows how the original loss can produce a biased endpoint distribution. Relatedly, Christensen et al. construct time-homogeneous denoising diffusions using Doob h-transforms with random horizons, but their construction is stochastic and score-based, rather than a deterministic transport derived from a Beckmann current.

Poisson Flow: The choice ν = 1 also yields a closed-form Coulomb transport related to Poisson Flow. In the BTM construction, the field is generated by the signed source–sink distribution µ0 − µ1, whereas PFGM places the data charges on an augmented hyperplane and uses the induced source distribution on a distant hemisphere.

Noise-unconditioned models: Kadkhodaie and Simoncelli provide a blind-denoiser sampler, while Sun et al. study removing noise conditioning. For linear FM, their population target matches our autonomous drift up to orientation, but their analysis bounds error relative to a conditioned sampler rather than proving an exact pushforward.

Optimal transport: The divergence condition is classical in OT. Standard flow matching realizes a transport without minimizing the Benamou–Brenier action; analogously, the FM instance of BTM gives a dynamical realization of a Beckmann flux without minimizing the associated cost.

Consistency models and one-step generation: One-step generation has been pursued through consistency models, rectified flow, Flow Map Matching, MeanFlow, and Shortcut Models. These methods start from an underlying time-dependent flow and distill or self-distill it into a one-step network. Drifting learns a one-step map directly, by evolving the network's pushforward during training, using an MMD distance to measure progress. BTM also learns the map directly, but exploits a different structural fact: the terminal map of the autonomous flow satisfies a stationary conservation equation along characteristics.

Standard flow matching constructs a stochastic interpolant:

Is = αsx0 + βsx1, x0 ∼ µ0, x1 ∼ µ1, s ∈ [0, 1],

where αs and βs satisfy boundary conditions α0 = β1 = 1, α1 = β0 = 0. The probability distribution µs of Is coincides with that of the solution to the probability flow ODE Ẏs = bs(Ys), Y0 ∼ µ0, provided the time-dependent drift is obtained by minimizing the standard FM loss over functions b̂s(x) that depend on s.

The paper considers minimizing the same averaged FM loss over time-independent functions b̂(x):

b = arg min Es∼U[0,1],x0,x1 b̂(Is) − İs2,

and considers the autonomous ODE Ẋt(x0) = b(Xt(x0)), X0(x0) = x0. The population minimizer is b(x) = Ex0,x1,s[İs Is = x].

Assumption 1 (Regular source and singular target, d ≥ 2): "The base distribution µ0(dx) = ρ0(x) dx has a smooth, strictly positive, rapidly decaying density; in particular, ρ0 has finite moments of every order. The target distribution µ1 is supported on a compact smooth embedded k-dimensional manifold M1 ⊂ Rd without boundary, where k ≤ d − 2, and has a smooth, strictly positive density ρ1 on M1 with respect to its natural volume measure vol M1."

Assumption 2 (Straight FM interpolant with controlled endpoint behavior): "The interpolant has straight-line geometry, Is = αsx0 + (1 − αs)x1, where x0 ∼ µ0 and x1 ∼ µ1 are independent, and α ∈ C1([0, 1]), α0 = 1, α1 = 0, and α̇s < 0 for 0 ≤ s 0 and γ ≥ 0."

"Under Assumptions 1–2, for µ0-almost every x0 ∈ Rd, there exists a finite time τ(x0) > 0 such that the solution of (6) is defined and remains in Rd M1 for 0 ≤ t < τ(x0), and T(x0):= lim Xt(x0) as t↑τ(x0) exists and belongs to M1. In addition, the resulting endpoint map T satisfies Tµ0 = µ1."

Proof sketch: Integrating the continuity equation from s = 0 to s = 1 gives ∇ · j = µ0 − µ1, where j = ∫01 bsµs ds is the time-averaged probability flux. The minimizer b satisfies j = νb, where ν = ∫01 µs ds > 0. Hence ∇ · (νb) = µ0 − µ1. Since ν > 0, the fields j and b have the same oriented flow lines. Away from M1, ∇ · j = ρ0 > 0, so volumes transported by j expand. The FM current carries no mass to infinity, and near M1 it points toward the target. Hence almost every trajectory ends on M1. For A ⊆ M1, flux balance equates source mass inside the basin with sink mass at its endpoint: µ0(B A) = µ1(A), giving Tµ0 = µ1.

"The role of the singular target is visible from the divergence equation (7): it provides the localized sink at which the autonomous trajectories can end. Codimension also matters. In codimension one, a current satisfying the same divergence equation may pass through the target rather than stop there, and counterexamples show that the source–sink balance alone is not sufficient."

"There are two distinct clocks. The schedule αs, with s ∈ [0, 1], governs the interpolant Is and the original time-dependent flow Ẏs = bs(Ys). The autonomous flow (6) instead runs on its own clock t. Its oriented streamlines, and hence the transport map T, depend only on the straight-line geometry ax0 + (1 − a)x1, not on the interpolant clock."

The behavior of αs near s = 1 determines the magnitude of the autonomous drift near the target:

  • For α̇1 0): b(x) ≍ r(x)(γ/(γ+1)) → 0 near M1, with finite hitting time.

For a fixed positive weight ν of finite mass, consider the affine class:

J ν:= j ∈ L2(ν−1 dx; Rd): ∇ · j = µ0 − µ1,

where the divergence equality is understood distributionally. The weighted Beckmann problem:

j ν:= arg min (1/2)∫ Rd j(x)2/ν(x) dx

has a unique solution whenever J ν is nonempty. Set b:= j ν/ν.

"Under Assumption 1, let ν ∈ C1(Rd M1) be strictly positive and satisfy 0 < ∫ Rd ν(x) dx < ∞. Assume that J ν is nonempty, let j ν be the minimizer in (9), and suppose that b:= j ν/ν belongs to C1(Rd M1). Suppose moreover that there exist r0, κ > 0 such that j ν(x) · ∇r(x) ≤ −κj ν(x) < 0 for 0 < r(x) < r0. Then, for µ0-almost every x0, there exists τ(x0) ∈ (0, ∞] such that the solution of (6) is defined and remains in Rd M1 for 0 ≤ t < τ(x0), and the unique limit T(x0):= lim Xt(x0) as t↑τ(x0) belongs to M1, where t ↑ τ means t → ∞ when τ = ∞. In addition, the resulting endpoint map T satisfies Tµ0 = µ1."

Let j FM be the FM current of Proposition 1 and let ν = ∫01 µs ds be its occupation weight. Then ∫ Rd ν dx = 1, j FM ∈ J ν, and j FM has finite weighted energy. If the autonomous FM regression is restricted to gradient fields b = ∇φ, its population loss is, up to an additive constant, (1/2)∫ Rd ν∇φ2 dx − ∫ Rd j FM · ∇φ dx + (1/2)E[İs2]. If this restricted problem admits a sufficiently regular minimizer, its first-order condition gives ∇ · (ν∇φ) = µ0 − µ1. Since ν∇φ is then feasible and is orthogonal in L2(ν−1 dx) to every finite-energy divergence-free perturbation, it coincides with the unique minimizer j ν.

"Under Assumption 1, let ν ≡ 1 and let b = ∇φ be the canonical Coulomb field satisfying ∆φ = µ0 − µ1, equivalently b(x) = (1/ω d)∫ Rd (x − y)/x − yd (µ0 − µ1)(dy), for x ∉ M1, where ω d is the surface area of the unit sphere in Rd. Then the first-hitting map of the autonomous flow Ẋt = b(Xt) transports µ0 to µ1, and its hitting time is finite for µ0-almost every initial point."

The finite-mass theorem does not include the Coulomb choice ν ≡ 1 because both relevant finiteness conditions fail: ∫ Rd ν dx = ∞, and the behavior j(x) ≍ r(x)(1−(d−k)) near M1 makes ∫ Rd j2 dx diverge. Nevertheless, it obeys the same transport conclusion because its neutral far field gives no loss at infinity and its singular part points uniformly toward M1.

The classical Beckmann transportation problem is:

B(µ0, µ1):= inf∇·j=µ0−µ1 ∫ Rd j(x) dx = W1(µ0, µ1).

The finite-mass currents constructed in this paper satisfy the same divergence constraint and are therefore competitors in (14), but we do not claim that they solve this L1 minimization problem.

Problem (9) minimizes a quadratic functional with the weight ν fixed. For every fixed current j, inf ν≥0, ∫ν dx=1 ∫ j2/ν dx = (∫ j dx)2, with equality for ν ∝ j. Thus optimizing both the weight and the current would recover the square of the classical Beckmann problem.

The structural analogy is:

  • Monge–Kantorovich optimal transport ↔ Benamou–Brenier (time-dependent flow)

  • Beckmann transportation ↔ Beckmann Transport Models (autonomous flow)

For the unconstrained FM current j FM, with b FM = j FM/ν and ∫ Rd ν dx = 1:

W12(µ0, µ1) ≤ (∫ Rd j FM dx)2 ≤ ∫ Rd j FM2/ν dx = ∫ Rd b FM2ν dx ≤ ∫01∫ Rd bs2 dµs ds.

Thus the FM current provides an upper bound on the value of the classical Beckmann problem, rather than solving it.

When τ(x0) < ∞, the flow is extended to all later times by freezing after arrival: Xt(x0):= T(x0) for t ≥ τ(x0). Under this convention, Xt inherits the semigroup property X t+s = Xt ∘ Xs for all s, t ≥ 0, with X0 = id.

"Let b be one of the autonomous drifts in Proposition 1, Theorem 1, or Proposition 2, and let Xt be the corresponding autonomous flow (6), with the freezing convention above. Viewed as a flow map of the initial condition x ∈ Rd, Xt(x) satisfies, wherever the spatial derivative exists and for almost every t, the Eulerian equation ∂tXt(x) = b(x)·∇Xt(x), X0(x) = x, for x ∉ M1, Xt(x) = x for x ∈ M1."

Proof: The semigroup identity gives X t+s(x) = Xt(Xs(x)). Differentiating in s at s = 0 yields ∂tXt(x) = b(x) · ∇Xt(x). For t τ(x), the left-hand side vanishes by freezing and the same semigroup argument gives b · ∇Xt = 0.

"Let T be the endpoint map of Proposition 1, Theorem 1, or Proposition 2. The endpoint is constant along µ0-almost every characteristic: T(Xt(x)) = T(x), 0 ≤ t < τ(x). After setting T(x) = x on M1, this gives, wherever T is differentiable, b(x) · ∇T(x) = 0 for x ∉ M1, T(x) = x for x ∈ M1."

The terminal-limit construction uniquely defines T up to a µ0-null set. "By contrast, if (18) is considered as a standalone stationary boundary-value problem, the boundary value must be understood as a trace along terminal characteristics; pointwise values on the lower-dimensional set M1 alone do not ensure uniqueness."

From the conservation equation (18), one could learn T by minimizing the squared residual E[b · ∇T2]. Proposition 3 provides a more direct motivation: the transport map is the terminal limit of the flow-map evolution, and a formal explicit-Euler step of (17) reads:

X(k+1)(x) = X(k)(x) + η b(x) · ∇X(k)(x).

The algorithm tested uses a single network on both sides and stops gradients through the right-hand side. In the FM setup, the regression identity b(x) = Es,x0,x1[İs Is = x] allows using the conditional sample İs · ∇T(Is).

"Under Assumptions 1–2, suppose that the terminal map T is differentiable almost everywhere with respect to the FM occupation measure and that the expectations below are finite. Then the population stop-gradient update vanishes at T for the objective:

L(T) = Es,x0,x1[T(Is) − sg(T(Is) + İs · ∇T(Is))2] + λ Ex1[T(x1) − x12]."

The theorem establishes "that the proposed learning rule is consistent with the true terminal map at population level. As is typical for stop-gradient objectives, it does not by itself prove uniqueness of the fixed point or convergence of neural-network optimization."

"Once the map is trained, sampling at inference is a single forward pass: x1 = T θ(x0) for x0 ∼ µ0. The semigroup property of the exact flow motivates testing whether repeated application of a partially trained network improves its approximation to the terminal map. Although a partially trained network is not guaranteed to equal Xt for any t, empirically its iterates sharpen samples."

Wang and Du proposed Equilibrium Matching via a loss that resembles (5) but pairs the linear interpolant Is = (1 − s)x0 + sx1 with a target cs(x1 − x0) scaled by a non-trivial schedule:

Es,x0,x1[b((1 − s)x0 + sx1) − cs(x1 − x0)2].

"This coincides with the loss in (5) only when cs ≡ 1, since İs = x1 − x0 for the stated linear interpolant. For other choices of cs, the regression target is not the velocity of that interpolant and therefore does not enforce the divergence condition (7). It will generally introduce bias and fail to transport µ0 to µ1."

"The resolution, which is what (5) implements, is to treat the schedule as part of the interpolant rather than as a separate weight on the target. To satisfy the desired tail behaviour, one may choose the straight interpolant of Assumption 2 with α̇1 = 0, equivalently İ1 = 0. The resulting autonomous drift extends continuously by zero to the support of µ1."

Algorithm 1 (Learning the autonomous drift b): Samples a minibatch B; for each k ∈ B, samples sk ∼ U([0, 1]), xk0 ∼ µ0, xk1 ∼ µ1; sets Ik = α skxk0 + β skxk1, İk = α̇ skxk0 + β̇ skxk1; updates θ ← θ − η∇ θ L, where L = (1/B)Σ k∈Bb θ(Ik) − İk2.

Algorithm 2 (Learning the transport map T with the base objective): Samples a minibatch B; for each k ∈ B, samples sk ∼ U([0, 1]), xk0 ∼ µ0, xk1 ∼ µ1; sets Ik = α skxk0 + β skxk1, İk = α̇ skxk0 + β̇ skxk1, T̂k = T θ(Ik) + İk · ∇T θ(Ik); updates θ ← θ − η∇ θ L, where L = (1/B)Σ k∈B[T θ(Ik) − sg(T̂k)2 + λT θ(xk1) − xk12].

The paper tests whether BTM, learned with the consistent loss (5), corrects the bias produced by the original EqM loss (21). They consider a two-dimensional example where µ0 = N(0, I2) and µ1 = Σⱼ=15 pⱼδ xⱼ, with deliberately unequal weights p = (0.30, 0.30, 0.15, 0.15, 0.10). For BTM, the five basin areas match the target weights to within measurement error (MAE = 0.005). For EqM, the two largest basins expand while the smallest nearly disappears (MAE = 0.102, a 20× error).

Using a two-dimensional spiral distribution embedded in d = 6 dimensions (µ0 = N(0, I6)) with Is = (1 − s)x0 + sx1, the paper compares generation with the autonomous flow associated with b against the learned one-step map T θ and its iterates. "The autonomous flow recovers the target sharply, while a one-shot map from a partially trained network is visibly diffuse around the spiral. Applying the trained network three times sharpens generation substantially and brings it close to the flow; after fuller training, one application matches the flow visually."

The paper trains the consistent loss (5) on class-conditional ImageNet 256 × 256, using the XL/2 architecture and the same training budget as Wang and Du [2025], with a self-stopping interpolant detailed in Appendix E. "The corrected model (EqM-XL/2) achieves FID = 1.87 versus FID = 1.90 for the uncorrected EqM-XL/2 under the same conditions. The gain is modest, as expected if the bias is primarily a weight-allocation error and the class distribution is nearly uniform."

The paper trains a one-step map using the stop-gradient loss (20) with an XL/2 architecture on an SD-VAE latent space, omitting time embeddings. For these ImageNet map experiments only, we additionally use adaptive weighting, gradient balancing, and a slightly noisy boundary anchor. Under the same budget as Wang and Du [2025], BTM achieves an FID of 17.58. The paper compares with iCT (FID 30.10), Shortcut Models (FID 10.60), SiT without CFG (FID 8.30), MeanFlow (FID 3.43), and SiT with CFG (FID 2.06). "While our unguided one-step result remains modest, the strongest competing one-step methods rely heavily on classifier-free guidance (CFG) for sample quality. Extending CFG to BTM remains an open direction for future work."

The paper introduces Beckmann Transport Models, a framework for generative modeling built on currents j = νb satisfying the divergence equation ∇ · j = µ0 − µ1. The transport property is proved separately for the finite FM current, for finite-mass weighted Beckmann minimizers with uniform terminal orientation, and for the canonical Coulomb current. The framework includes a Coulomb construction closely related to Poisson Flow and a flow-matching-consistent correction of Equilibrium Matching.

"The exact transport statements are population-level results. As in standard diffusion and flow-matching theory, implementations based on finite data and neural networks approximate the population field or map; the ImageNet experiments evaluate the proposed constructions in this practical regime rather than asserting finite-sample exactness."

Open directions: One could optimize jointly over admissible pairs (ν, b) subject to ∇ · (νb) = µ0 − µ1, while enforcing (23) and (24). The framework may also be relevant to text generation via embedding on the probability simplex: text tokens correspond to vertices of the simplex (atomic distributions), so the singular-target assumption is satisfied naturally. Finally, the empirical gains from iterated inference in Section 2.6 suggest a trade-off between training-time and inference-time compute that has not yet been characterized theoretically.

The appendix provides precise regularity conditions: "rapidly decaying means that ρ0 is a Schwartz density. Thus ρ0 ∈ C∞(Rd) is strictly positive and, for every multi-index β and integer N ≥ 0, sup (1 + x) N ∂ βρ0(x) < ∞." It is enough for M1 to be a compact embedded C3 submanifold without boundary and for its positive density ρ1 to belong to C1(M1). The natural volume measure is vol M1 = Hk ↾ M1.

Theorem 4 (Autonomous transport from flux balance): "Assume the regularity and measure conditions of Assumption 1, except that here it is enough that M1 have positive codimension, k ≤ d − 1. Define r(x):= dist(x, M1), and let ν, b ∈ C1(Rd M1) with ν > 0, set j:= νb, and assume that j ∈ L1 loc(Rd; Rd). Suppose that, for every ζ ∈ C c∞(Rd), ∫ Rd j · ∇ζ dx = −∫ Rd ζρ0 dx + ∫ M1 ζρ1 dvol M1, and that lim R→∞ (1/R)∫ R 0 and an r0 > 0 smaller than the tubular-neighborhood radius of M1 such that j(x) · ∇r(x) ≤ −κj(x) < 0 for 0 < r(x) < r0. Then, for µ0-almost every initial point, the forward solution of Ẋt = b(Xt) has a unique endpoint T(x) ∈ M1. Moreover, Tµ0 = µ1."

The proof uses a flow-tube identity, rules out escape to infinity and trapping in compact sets away from M1, establishes the endpoint distribution via flux balance, and shows the autonomous travel time is τ(x) = ∫ Γ dl/b.

The proof of Proposition 2 shows the distributional identity ∇ · (x/(ω dxd)) = δ0, establishes the far-field estimate b(x) = O(x−d) as x → ∞, and verifies the terminal direction near M1 with b(p + rξ) = −c d,kρ1(p)r 1−qξ + o(r 1−q), where c d,k:= (1/ω d)∫ Rk dw/(1 + w2) d/2 > 0.

For the standard Gaussian base, the canonical Coulomb field is explicit:

b(x) = (1/ω d)∫ M1 (x − y)/x − yd µ1(dy) − (x/xd)(γ(d/2, x2/2)/Γ(d/2)).

Changing the clock αs without changing the path leaves the current j unchanged. The occupation density does depend on the clock: ν(x) = ∫01 ρ̃a(x)/(−α̇ α−1(a)) da. "Consequently a clock change modifies b = j/ν and hence the speed of the autonomous flow. Because ν > 0, it does not change its flow lines." Changing the path geometry f changes the current and may change both flow lines and endpoints.

For a breakpoint s c ∈ (0, 1), the schedule is:

αs = 1 − 2s/(1 + s c) for 0 ≤ s ≤ s c,

αs = (1 − s)2/(1 − s c)2 for s c < s ≤ 1.

"The two branches have the same value and derivative at s c, so α ∈ C1([0, 1]). The schedule is linear before s c and quadratic near the endpoint. Hence İ1 = 0 and, under the assumptions of Proposition 1, b(x) ≍ dist(x, M1) 1/2 near M1. The hitting time remains finite because ∫0 ϵ r−1/2 dr < ∞."

BTM and the Drifting framework of Deng et al. [2026] both use a fixed-point viewpoint, but construct their dynamics differently. Drifting learns a McKean–Vlasov dynamics with kernel K:

ẋt = E[K(xt, Xt)Xt]/E[K(xt, Xt)] − E[K(xt, X∗)X∗]/E[K(xt, X∗)].

Kernel performance can be sensitive to scale and geometry in high dimensions. The experiments show Drifting is sensitive to the kernel choice, whereas BTM trains stably with both interpolants shown.

  • Bias grows monotonically with κ: At κ = 0 the two losses are algebraically identical and both give MAE ≈ 0.007. As κ increases, the EqM loss degrades monotonically (by up to a factor of 25 at κ = 0.9), while the consistent FM loss remains nearly flat.

  • Convergence speed vs. correctness: For αs = (1 − s) a, both a = 1 and a = 2 achieve accurate mass allocation (MAE ≈ 0.005). They differ mainly in speed: 99% of particles freeze by autonomous time t ≈ 4 for a = 1 and by t ≈ 6 for a = 2.

  • Training-free Coulomb transport: Mini-batch estimates of the Coulomb drift transport a Gaussian cloud toward a Swiss-roll target in d = 5 without a neural network or training loop.

MNIST digit generation: We approximate T with a standard 23-million-parameter diffusion U-Net after removing its time conditioning. Repeated application removes some visual artifacts and sharpens the digit boundaries.

Latent ImageNet 256 × 256 generation: Uses the SiT architecture without time embeddings in the SD-VAE latent space. Stabilization techniques include:

  • Adaptive weighting: w = 1/((∥∆∥22 + c) p) with p = 1 and c = 0.01.

  • Gradient balancing: λ = ∥∇ L L boundary∥/∥∇ L L transport∥.

  • ImageNet-only noisy boundary anchor: L anchor(T θ) = E s,x0,x1[1 s>0.8T θ(Is) − x12].

The BTM B/2 ablation model achieves FID 62.14 at 80 epochs (vs. MeanFlow B/2 at 61.06). The full BTM XL/2 model achieves FID 17.58 at 1280 epochs. Iterated application improves FID: 17.58 (1 NFE), 17.04 (5 NFE), 16.53 (10 NFE).

Improvements for AI systems

Based on the paper, here are specific improvements I can implement in AI systems, along with the resulting capabilities:

Improvement: Replace time-conditioned diffusion/flow-matching models with time-independent velocity fields that satisfy the divergence equation ∇·(νb) = μ0 − μ1. This eliminates the need for time embeddings in neural networks and simplifies the ODE solver.

Capability: A generative model that:

  • Uses a single static vector field instead of a time-dependent one

  • Requires no time-conditioning in the network architecture

  • Produces samples by integrating a simple autonomous ODE until hitting the target manifold

  • Achieves exact transport to the target distribution when the target is supported on a codimension-≥2 manifold

These improvements collectively enable: (1) simpler architectures without time conditioning, (2) one-step generation with optional iterative refinement, (3) exact mass preservation for singular targets, (4) training-free baselines for validation, and (5) a principled framework for discrete output generation.

Abstract

We propose an instantiation of flow matching that relies on a time-independent velocity field (an autonomous flow) to exactly map between two distributions, so long as the target is singular, i.e. supported on a lower-dimensional data manifold. We also show that the one-step generative map associated with this flow is the unique solution of a simple conservation equation, which can be used to learn the map directly from samples. These autonomous flows and maps give a dynamical meaning to the flux constraint of Beckmann's transportation problem. Their construction provides a unifying framework that recovers, for instance, the closed-form Poisson-flow generative model and equilibrium matching with a quadratic flow-matching regression loss. We illustrate how this theory corrects inconsistencies in existing methods and demonstrate the effectiveness of the autonomous flow and the one-step map on ImageNet 256x256.

Sources

Related papers