Stability of the Monge Map in Semi-Dual Optimal Transport

arXiv:2605.05569 · math.OC, cs.LG · Submitted 2026-05-07 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Stability of the Monge Map in Semi-Dual Optimal Transport".

Jane: The paper was written by Anton Selitskiy and David Millard from University of Rochester and Rochester Institute of Technology and Department of Electrical and Computer Engineering, University of Rochester and Department of Mechanical Engineering, Rochester Institute of Technology.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: We’re looking at this paper, "Stability of the Monge Map in Semi-Dual Optimal Transport," and it's making a big claim about how the mathematical structure of optimal transport itself behaves.

Jane: It essentially says that the way we usually define these problems, using a semi-dual approach, has a specific structural weakness called a degenerate saddle-point.

Lu: This means that when we are at the perfect transport solution, what we are trying to optimize—the objective function—becomes completely independent of one of the key variables: the potential.

Meng: That independence is really problematic for standard optimization algorithms because they expect every variable to play a role in defining the best outcome.

Lalam: The implication here is that simply finding a local minimum isn't enough if that optimal point has this inherent flatness, which could lead to spurious or unstable solutions.

Tom: It’s not just about the math, Jane; it’s about how we train these AI systems based on those mathematical structures.

Jane: Exactly, Tom. The paper highlights that this structural degeneracy is a major factor in "Stability of the Monge Map in Semi-Dual Optimal Transport."

Lu: It suggests that simply finding an optimal potential function might not be enough to guarantee a stable training process for the AI model.

Meng: So, we need to figure out how to handle scenarios where the math itself is inherently weak at certain points of learning.

Lalam: By identifying this degeneracy, they are showing us where our current understanding of stability needs refinement in machine learning models.

Summary: Tom: Let's talk about what the paper summarizes, which is that despite this degeneracy, there's a way to measure convergence without needing the dual potential to be perfect.

Jane: The authors provide a specific formula—a key estimate—that allows us to track how close our transport map gets to the true optimal map.

Lu: This metric, as shown in "Stability of the Monge Map in Semi-Dual Optimal Transport, " gives us a quantifiable measure of error that doesn's rely on the assumption that we have found a perfect potential.

Meng: That error estimate is crucial because it provides a concrete way to judge performance even when the underlying mathematical assumptions about optimal solutions aren't met.

Lalam: It’s like having a reliable dashboard for our AI model, telling us how far off we are from perfection, even if the model isn't fully optimized yet.

Tom: It gives us a way to quantify the success of the transport map even when we don't have that optimal potential function.

Jane: That’s right, Tom. This move away from requiring perfect dual potential is a major conceptual shift in "Stability of the Monge Map in Semi-Dual Optimal Transport."

Lu: It implies that we can measure progress based solely on the performance of the map itself, which is a much more robust way to track learning.

Meng: If this estimate holds, it means we can design better stopping criteria for our AI training loops without getting misled by a non-optimal potential.

Lalam: The ability to quantify convergence independently is a huge step toward ensuring reliable and predictable behavior in generative AI.

Improvements: Tom: Building on the idea that optimal solutions are hard to define, the paper suggests a practical explanation for why some AI algorithms struggle with convergence.

Jane: They found that in practice, we usually need many more updates for the transport map than we need for the potential function.

Lu: This observation points toward interpreting the problem as a continuous-time dynamical system, showing that two variables are updating at different speeds.

Meng: That concept of two timescales is what's really practical; it tells us exactly where to focus our engineering efforts in "Stability of the Monge Map in Semi-Dual Optimal Transport."

Lalam: The idea that the potential evolves slower than the map suggests a natural, steady progression toward stability in AI training.

Tom: So, if we know this dynamic is real, we can design an algorithm that handles it without crashing or stalling.

Jane: It’s about recognizing that "Stability of the Monge Map in Semi-Dual Optimal Transport" describes a process where the map is the fast mover and the potential is more deliberate.

Lu: This insight allows us to move beyond just theoretical convergence and actually implement a practical, robust training schedule for AI.

Meng: It gives us clear guidance on how to structure our optimization loops so that we aren't waiting on a slow dual variable when the primary map is ready to advance.

Lalam: By formalizing this two-timescale relationship, we are moving toward a more sophisticated and efficient way of building generative models in AI.

Conclusion: Tom: So, we've covered how "Stability of the Monge Map in Semi-Dual Optimal Transport" reveals fundamental flaws in our assumptions about optimal solutions.

Jane: We've seen that this degeneracy is not only a theoretical curiosity but a practical obstacle to reliably measuring convergence in AI.

Lu: The idea that the potential isn't always necessary for achieving stability is a monumental shift for theoretical computer science.

Meng: We can now design more robust training pipelines because we understand exactly how the map and potential operate at different speeds.

Lalam: This work allows us to move toward a more mature, predictable culture in AI research, ensuring that our models perform as intended.

Tom: It's a paper that provides deep structural understanding of optimal transport and its practical implications for reliability.

Jane: We really hope this work by the authors sets a new standard for how we evaluate and improve generative AI.

Lu: I think the mathematical rigor applied here is exactly what' needed to bring these complex problems into reality.

Meng: It offers clear guidance on implementation, making "Stability of the Monge Map in Semi-Dual Optimal Transport" an incredibly useful tool for engineers too.

Lalam: By understanding this stability, we are paving the way for a more reliable future for AI applications globally.

University of Rochester · Rochester Institute of Technology · Department of Electrical and Computer Engineering, University of Rochester · Department of Mechanical Engineering, Rochester Institute of Technology

math.OC, cs.LG

Submitted: 2026-05-07

Updated: 2026-09-07

Importance score: 92/100

The gist: This paper investigates the mathematical structure of the semi-dual formulation of optimal transport (OT), specifically addressing why numerical algorithms for learning transport maps often exhibit

Key concepts

Degenerate Saddle-Point
A structural weakness in optimal transport where, at the perfect solution, what is being optimized becomes completely independent of a key variable: the potential. This independence makes standard optimization algorithms problematic.
Convergence Metric
The authors provide a specific formula or key estimate that allows researchers to track how close the transport map gets to its true optimal state. This quantifiable measure of error does not rely on the assumption that a perfect dual potential has been found.
Two Timescales
This concept describes an observed dynamic where, in practice, the transport map requires significantly more updates than the dual potential function. This indicates that these two variables are updating at different speeds during AI training.

Terminology

Summary

This paper investigates the mathematical structure of the semi-dual formulation of optimal transport (OT), specifically addressing why numerical algorithms for learning transport maps often exhibit different convergence behaviors than those for dual potentials. By identifying a degenerate saddle-point structure in neural optimal transport (NOT), the authors provide a theoretical framework to explain training instabilities and establish new convergence guarantees for Monge maps that do not rely on the optimality of the dual potential.

The core discovery

The authors demonstrate that the semi-dual optimal transport problem possesses a degenerate saddle-point structure, meaning that once an optimal transport map is reached, the objective functional becomes independent of the potential. This phenomenon implies that while a transport map may converge to its optimal state, the dual potential may continue to change without affecting the objective. The paper notes that this closely parallels instability phenomena in generative adversarial networks (GANs), where a discriminator might change without impacting the generator's optimality.

This degeneracy has significant implications for existing literature:

** Many convergence estimates implicitly assume that the transport map exists and that the optimization problem is well-posed. 1 2 3 4 5 6 7 **

** Several works rely on properties derived under the assumption that the potential is optimal, which is neither necessary nor enforced by the optimization dynamics once the transport map is close to optimal. **

Convergence and stability

The paper provides a new estimate for the transport map that is independent of the potential, expressed as:

T − T⋆2 L2 ≲ E[c(x, T(x))] − W2(µ, ν)2 + d KR(Tµ, ν).

This result allows for the assessment of convergence through observable quantities rather than exact Kantorovich distances. The authors argue that because the constraint (the marginal condition) is only enforced approximately in practice, convergence must be assessed by simultaneous improvement in both the transport cost and the discrepancy between distributions. They further interpret the training dynamics as a continuous-time two-timescale dynamical system, where:

  1. The primal variable (the transport map) evolves on a faster timescale to stabilize the system.

  2. The dual variable (the potential) evolves more slowly to enforce the marginal constraint.

Connection to generative modeling

The research situates NOT within the broader landscape of generative models, noting that GANs, VAEs, and diffusion models can all be interpreted as approximations of transport processes. The authors classify existing static NOT solvers into three distinct groups:

** Group I: min–max optimization over the transport map T and the potential φ (or ψ). **

** Group II: min–max optimization over φ (or ψ) with the transport map defined via gradients (e.g., T = ∇φ). **

** Group III: regularized or unbalanced formulations. **

By analyzing these formulations, the paper explains that spurious solutions in the semi-dual formulation arise from the degeneracy. Near optimal maps, the dual potential becomes non-identifiable, meaning different potentials can yield distinct gradients, potentially leading optimization to stationary points that do not correspond to the true Wasserstein distance. Finally, they conclude that NOT can be viewed as a constrained optimization problem similar to unbalanced OT, where the marginal constraint is relaxed via penalization.of the transport map.

Improvements for AI systems

To improve AI systems using the findings from Stability of the Monge Map in Semi-Dual Optimal Transport, I would implement the following specific architectural and algorithmic changes:

  1. Implement a Two-Timescale Asymmetric Update schedule for Generative Adversarial Networks (GANs) and Neural Optimal Transport (NOT) solvers.

  2. Transition from standard minimax objectives to Constrained Optimization formulations that treat the marginal distribution constraint as a soft penalty term rather than relying on the dual potential's optimality.

These improvements would enable an AI system to:

  1. Decouple map convergence from potential convergence, allowing for highly accurate transport maps (e.g., for high-fidelity image synthesis or density estimation) even when the discriminator/potential has not yet reached its optimal state, significantly reducing training time and computational waste.

  2. Eliminate Gradient Deviation and spurious solutions in generative modeling by acknowledging the inherent degeneracy of the saddle-point structure, preventing the model from stalling at non-optimal stationary points.

  3. Achieve more stable training in high-dimensional spaces by using multiple primal updates (transport map) for every single dual update (potential), effectively acting as a stabilizer that prevents the unboundedness issues typically seen when the marginal constraint is not yet satisfied.

  4. Provide reliable convergence diagnostics by monitoring the simultaneous reduction of both transport cost and marginal discrepancy, rather than relying on potentially misleading objective values that may appear converged due to functional flatness.

Sources

Related papers