A primer on optimal transport for causal inference with observational data
Florian Gunsilius
Department of Economics, Emory University
stat.ME, cs.AI, econ.EM
Submitted: 2026-08-07
Comments: Updated section 4.2
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 81/100
The gist: This review, "A primer on optimal transport for causal inference with observational data," aims to "unify the language and notation between different areas of statistics, mathematics, and
Terminology
Summary
This review, A primer on optimal transport for causal inference with observational data,
aims to unify the language and notation between different areas of statistics, mathematics, and econometrics
by demonstrating that optimal transport is not just a set of potential tools, but actually builds the foundation of model assumptions
for identifying causal effects.
Foundations of Optimal Transport and Structural Models
The paper defines the fundamental Monge problem as finding an optimal map T: X to Y that preserves the mass between P and Q and minimizes the overall cost of transporting P onto Q.
This is relaxed via the Kantorovich problem, which seeks an optimal coupling gamma between P and Q.
In the case of squared Euclidean distance, the value function is the 2-Wasserstein distance,
and Brenier’s theorem implies the optimal map is the gradient of a convex function.
In structural models of the form Y = g(X, U), where U represents unobservable heterogeneity, the paper establishes a connection to optimal transport in exogenous settings (X U). It states that the identification argument... implies that g(x, times) is the monotone rearrangement between the unobservable F U and the observable F YX=x for P X-almost every x.
This monotone rearrangement is the unique measure- and order preserving transformation
and serves as the identifying assumption.
Instrumental Variable (IV) Models
When X is endogenous (X U), the paper explores several approaches:
-
Control Variables: This approach mimics conditioning on unobservables. The paper notes that
the control variable approach... [uses] the monotone rearrangement... to identify average structural effects.
-
Fixed-Point Iterations: For binary instruments, the paper describes a method to identify heterogeneous effects by alternating between
vertical
maps (changing the instrument realization) andhorizontal
maps (changing quantiles). This processeither converges or diverges
to a fixed point, allowing the identification of the mechanism g(x, V). In multivariate settings, this involves thedynamics of Brenier maps
and theinverse Brenier rank map.
-
Partial Identification: To obtain bounds on average structural effects without strong structural assumptions, the paper describes a
generalized optimal transport problem on path spaces, similar to the stochastic optimal transport problem.
This approach seeks a measure on thejoint path space of the processes
that is consistent with observed marginals. -
Distributionally Robust Optimization (DRO): The paper discusses using
Wasserstein balls
to define regions of possible distributions, noting that theprimal DRO problem... admits a dual problem that often takes the form of a standard constrained prediction problem.
Difference-in-Differences (DiD) and Synthetic Controls
-
Nonlinear DiD: The
changes-in-changes estimator
is identified as beingbased on the monotone rearrangement.
It works by attempting totransplant the 'natural trend' d —that is, the change in the outcome in the case where no treatment is administered—to the outcome P Y T,0 of the treatment group before treatment.
The paper also discusses a multivariate extension usingcyclically comonotone production functions
and Brenier maps. -
Systems View: Beyond individual heterogeneity, the paper introduces
distributional synthetic controls,
which focus onhow the entire distribution of the outcome changes under different counterfactual states.
This method isbased on optimal transportation, in particular barycenters in Wasserstein space.
-
Synthetic Controls: While classic synthetic controls use a
convex combination of such control units
to replicate pre-treatment trends, the distributional version estimates thecounterfactual quantile function F Y 0t,N-1(q)... by an optimally weighted average of the control quantile functions.
Matching and Unbalanced Optimal Transport
The paper identifies a limitation in classic matching: the mass-preserving constraint in the Monge-Kantorovich problem implies that all individuals in both groups have to be matched,
which introduces excessive bias into the estimator
if the overlap of supports is imperfect. To resolve this, the paper proposes unbalanced optimal transportation... which relaxes the measure-preservation constraint, hence allowing for individuals to remain unmatched.
This approach uses phi-divergence terms to allow for the creation and destruction of mass,
providing partial optimal matches
where individuals without close matches are automatically discarded.
Improvements for AI systems
1. Structural Causal Disentanglement Layers
By integrating monotone rearrangement constraints into neural network architectures, AI systems can explicitly separate observable features (X) from unobservable heterogeneity (U). This allows the system to learn the true structural mechanism g(x, u) rather than mere correlations, enabling accurate what-if
simulations in environments where hidden confounders are present.
2. Wasserstein-DRO Policy Optimization
By implementing Distributionally Robust Optimization (DRO) using Wasserstein balls within Reinforcement Learning (RL), agents can be trained to optimize for the worst-case distribution within a specified distance of the training data. This makes AI agents significantly more robust to environmental shifts, distribution drift, and out-of-distribution (OOD) scenarios in real-world deployment.
3. Counterfactual Distributional Forecasting
Instead of predicting single-point estimates for the impact of an intervention, AI systems can use Wasserstein barycenters to predict the entire counterfactual probability distribution. This allows decision-support systems to perform advanced risk assessment, such as predicting how an intervention might affect the variance, skewness, or tail risks (extreme events) of an outcome.
4. Unbalanced Domain Adaptation Modules
By utilizing unbalanced optimal transport with phi-divergence terms in transfer learning, AI models can perform domain adaptation even when the source and target domains have non-overlapping supports. The system can discard
non-transferable data points (mass destruction) rather than forcing incorrect alignments, drastically reducing bias in cross-domain tasks like synthetic-to-real image translation or cross-lingual transfer.
5. Heterogeneous Effect Identification via Fixed-Point Brenier Maps
In multi-agent or personalized AI systems, incorporating fixed-point iterations of Brenier maps allows the system to identify how specific instruments
(e.g., a specific prompt, a policy change, or a nudge) affect different sub-populations of users differently. This enables the AI to move beyond average treatment effects
to provide highly personalized, heterogeneous response modeling.
Abstract
The theory of optimal transportation has developed into a powerful and elegant framework for comparing probability distributions, with wide-ranging applications in all areas of science. The fundamental idea of analyzing probabilities by comparing their underlying state space naturally aligns with the core idea of causal inference, where understanding and quantifying counterfactual states is paramount. Despite this intuitive connection, explicit research at the intersection of optimal transport and causal inference is only beginning to develop. Yet, many foundational models in causal inference have implicitly relied on optimal transport principles for decades, without recognizing the underlying connection. Therefore, the goal of this review is to offer an introduction to the surprisingly deep existing connections between optimal transport and the identification of causal effects with observational data -- where optimal transport is not just a set of potential tools, but actually builds the foundation of model assumptions. As a result, this review is intended to unify the language and notation between different areas of statistics, mathematics, and econometrics, by pointing out these existing connections, and to explore novel problems and directions for future work in both areas derived from this realization.
Sources
- Conservative Inference for Counterfactuals
- The Balancing Act in Causal Inference
- Optimal transport and Wasserstein distances for causal models
- Lorenz map, inequality ordering and curves based on multidimensional rearrangements
- A path-sampling method to partially identify causal effects in instrumental variable models
- Matching for causal effects via multimarginal unbalanced optimal transport
- Distributionally Robust Instrumental Variables Estimation
- Return to Office and the Tenure Distribution
- Asymptotic Properties of the Distributional Synthetic Controls
Related papers
- Doubly robust inference via calibration
- Bayesian Empirical Bayes: Simultaneous Inference from Probabilistic Symmetries
- Flexible Nonparametric Inference for Causal Effects under the Front-Door Model
- Deployment of AI-Assisted Interventions: Capacity Constraints and Noisy Compliance
- A Survey on Archetypal Analysis
- Dynamic Spatial Bayesian Machine Learning Model: Applications to Intergenerational Economic Mobility and Geographic Income Inequality in the United States