A primer on optimal transport for causal inference with observational data

arXiv:2503.07811 · stat.ME, cs.AI, econ.EM · Submitted 2026-08-07 · Read on arXiv

Florian Gunsilius

Department of Economics, Emory University

stat.ME, cs.AI, econ.EM

Submitted: 2026-08-07

Comments: Updated section 4.2

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 81/100

The gist: This review, "A primer on optimal transport for causal inference with observational data," aims to "unify the language and notation between different areas of statistics, mathematics, and

Terminology

Summary

This review, A primer on optimal transport for causal inference with observational data, aims to unify the language and notation between different areas of statistics, mathematics, and econometrics by demonstrating that optimal transport is not just a set of potential tools, but actually builds the foundation of model assumptions for identifying causal effects.

Foundations of Optimal Transport and Structural Models

The paper defines the fundamental Monge problem as finding an optimal map T: X to Y that preserves the mass between P and Q and minimizes the overall cost of transporting P onto Q. This is relaxed via the Kantorovich problem, which seeks an optimal coupling gamma between P and Q. In the case of squared Euclidean distance, the value function is the 2-Wasserstein distance, and Brenier’s theorem implies the optimal map is the gradient of a convex function.

In structural models of the form Y = g(X, U), where U represents unobservable heterogeneity, the paper establishes a connection to optimal transport in exogenous settings (X U). It states that the identification argument... implies that g(x, times) is the monotone rearrangement between the unobservable F U and the observable F YX=x for P X-almost every x. This monotone rearrangement is the unique measure- and order preserving transformation and serves as the identifying assumption.

Instrumental Variable (IV) Models

When X is endogenous (X U), the paper explores several approaches:

  • Control Variables: This approach mimics conditioning on unobservables. The paper notes that the control variable approach... [uses] the monotone rearrangement... to identify average structural effects.

  • Fixed-Point Iterations: For binary instruments, the paper describes a method to identify heterogeneous effects by alternating between vertical maps (changing the instrument realization) and horizontal maps (changing quantiles). This process either converges or diverges to a fixed point, allowing the identification of the mechanism g(x, V). In multivariate settings, this involves the dynamics of Brenier maps and the inverse Brenier rank map.

  • Partial Identification: To obtain bounds on average structural effects without strong structural assumptions, the paper describes a generalized optimal transport problem on path spaces, similar to the stochastic optimal transport problem. This approach seeks a measure on the joint path space of the processes that is consistent with observed marginals.

  • Distributionally Robust Optimization (DRO): The paper discusses using Wasserstein balls to define regions of possible distributions, noting that the primal DRO problem... admits a dual problem that often takes the form of a standard constrained prediction problem.

Difference-in-Differences (DiD) and Synthetic Controls

  • Nonlinear DiD: The changes-in-changes estimator is identified as being based on the monotone rearrangement. It works by attempting to transplant the 'natural trend' d —that is, the change in the outcome in the case where no treatment is administered—to the outcome P Y T,0 of the treatment group before treatment. The paper also discusses a multivariate extension using cyclically comonotone production functions and Brenier maps.

  • Systems View: Beyond individual heterogeneity, the paper introduces distributional synthetic controls, which focus on how the entire distribution of the outcome changes under different counterfactual states. This method is based on optimal transportation, in particular barycenters in Wasserstein space.

  • Synthetic Controls: While classic synthetic controls use a convex combination of such control units to replicate pre-treatment trends, the distributional version estimates the counterfactual quantile function F Y 0t,N-1(q)... by an optimally weighted average of the control quantile functions.

Matching and Unbalanced Optimal Transport

The paper identifies a limitation in classic matching: the mass-preserving constraint in the Monge-Kantorovich problem implies that all individuals in both groups have to be matched, which introduces excessive bias into the estimator if the overlap of supports is imperfect. To resolve this, the paper proposes unbalanced optimal transportation... which relaxes the measure-preservation constraint, hence allowing for individuals to remain unmatched. This approach uses phi-divergence terms to allow for the creation and destruction of mass, providing partial optimal matches where individuals without close matches are automatically discarded.

Improvements for AI systems

1. Structural Causal Disentanglement Layers

By integrating monotone rearrangement constraints into neural network architectures, AI systems can explicitly separate observable features (X) from unobservable heterogeneity (U). This allows the system to learn the true structural mechanism g(x, u) rather than mere correlations, enabling accurate what-if simulations in environments where hidden confounders are present.

2. Wasserstein-DRO Policy Optimization

By implementing Distributionally Robust Optimization (DRO) using Wasserstein balls within Reinforcement Learning (RL), agents can be trained to optimize for the worst-case distribution within a specified distance of the training data. This makes AI agents significantly more robust to environmental shifts, distribution drift, and out-of-distribution (OOD) scenarios in real-world deployment.

3. Counterfactual Distributional Forecasting

Instead of predicting single-point estimates for the impact of an intervention, AI systems can use Wasserstein barycenters to predict the entire counterfactual probability distribution. This allows decision-support systems to perform advanced risk assessment, such as predicting how an intervention might affect the variance, skewness, or tail risks (extreme events) of an outcome.

4. Unbalanced Domain Adaptation Modules

By utilizing unbalanced optimal transport with phi-divergence terms in transfer learning, AI models can perform domain adaptation even when the source and target domains have non-overlapping supports. The system can discard non-transferable data points (mass destruction) rather than forcing incorrect alignments, drastically reducing bias in cross-domain tasks like synthetic-to-real image translation or cross-lingual transfer.

5. Heterogeneous Effect Identification via Fixed-Point Brenier Maps

In multi-agent or personalized AI systems, incorporating fixed-point iterations of Brenier maps allows the system to identify how specific instruments (e.g., a specific prompt, a policy change, or a nudge) affect different sub-populations of users differently. This enables the AI to move beyond average treatment effects to provide highly personalized, heterogeneous response modeling.

Abstract

The theory of optimal transportation has developed into a powerful and elegant framework for comparing probability distributions, with wide-ranging applications in all areas of science. The fundamental idea of analyzing probabilities by comparing their underlying state space naturally aligns with the core idea of causal inference, where understanding and quantifying counterfactual states is paramount. Despite this intuitive connection, explicit research at the intersection of optimal transport and causal inference is only beginning to develop. Yet, many foundational models in causal inference have implicitly relied on optimal transport principles for decades, without recognizing the underlying connection. Therefore, the goal of this review is to offer an introduction to the surprisingly deep existing connections between optimal transport and the identification of causal effects with observational data -- where optimal transport is not just a set of potential tools, but actually builds the foundation of model assumptions. As a result, this review is intended to unify the language and notation between different areas of statistics, mathematics, and econometrics, by pointing out these existing connections, and to explore novel problems and directions for future work in both areas derived from this realization.

Sources

Related papers