Huber-Wasserstein barycenters for robust distribution-valued data
Carlos Cardoso-Perelló, Alberto González-Sanz
Columbia University
stat.ME, math.PR, stat.ML
Submitted: 2026-08-13
Updated: 2026-08-14
Comments: 75 pages, 4 figures
Code: https://github.com/carlosaccp/huber_ot
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 75/100
Terminology
Summary
arXiv: 2608.13131v1 [stat.ME] 13 Aug 2026
The paper proposes a robust barycenter for distribution-valued data by incorporating the Huber loss directly into the optimal transport cost. Unlike metric-space Huber means that apply the Huber loss to the Wasserstein distance after optimization, this construction acts on individual transport displacements, preserving quadratic behavior locally while limiting the influence of large displacements. The resulting Huber–Wasserstein barycenters form a natural interpolation between Wasserstein means and L1-type Wasserstein medians. The paper establishes analytical and statistical foundations including regularity and uniqueness properties of dual potentials, existence of optimal transport maps, stability as the Huber parameter varies, existence and characterization results for the barycenter problem, consistency of empirical plug-in estimators, and a finite-sample breakdown point essentially equal to 1/2. In dimension one, the paper derives the pointwise influence function and asymptotic distribution, quantifies the robustness–efficiency tradeoff, and shows that displacement-wise Huberization can retain first-order information lost by distance-based Huberization under localized shape contamination.
Modern statistical problems increasingly involve data objects that do not naturally belong to Euclidean spaces, particularly distribution-valued data where each observation is itself a probability measure. Examples include mortality and age-at-death distributions, regional income and house-price distributions, neuroimaging intensity distributions, particle-size distributions in geoscience, and collections of histograms or empirical measures.
The standard approach for defining a center is the Fréchet mean using the p-Wasserstein distance. For p = 2, the Wasserstein barycenter is defined as:
m(P):= argmin ν∈P2(Rd) ∫ W22(µ, ν) dP(µ)
While Wasserstein barycenters have been extensively studied, they can be highly sensitive to outlying distributions. The paper notes several existing robust approaches:
-
Impartial trimming (Álvarez-Esteban et al., 2008, 2018): trimming mass within distributions or entire distribution-valued observations
-
Wasserstein medians (Carlier et al., 2024; You et al., 2025)
-
Metric-space Huber means (Lee and Jung, 2026): applying the Huber loss outside the optimal transport problem
The metric-space Huber mean is defined as:
m′ H,c(P):= argmin ν∈P2(Rd) ∫ ρ c(W2(µ, ν)) dP(µ)
where ρ c is the Huber loss with parameter c > 0:
-
ρ c(t) = ½t2 if t < c
-
ρ c(t) = ct − c2/2 if t ≥ c
However, this approach has difficulties because the map ν ↦ ρ c(W2(µ, ν)) does not inherit the same convexity structure as the squared-Wasserstein objective.
The paper proposes a fundamentally different robustification strategy: instead of applying the Huber loss to the Wasserstein distance itself, the Huber loss is incorporated at the level of the optimal transport cost. Specifically:
m c(P):= argmin ν∈P1(Rd) Γ c,P(ν)
where
Γ c,P(ν):= ∫ [T ρ c(ν, µ) − c∫∥x∥ dµ(x)] dP(µ)
and
T ρ c(µ, ν):= inf π∈Π(µ,ν) ∫ ρ c(∥x − y∥) dπ(x, y)
is the optimal transport cost associated with the Huber loss. The centering term c∫∥x∥ dµ(x) ensures that Γ c,P(ν) is finite for every P ∈ P(P1(Rd)) without requiring integrability conditions on the first moments.
This construction differs fundamentally from the metric-space Huber mean by modifying the transportation cost itself, so the Huber loss acts locally on individual transport displacements. The objective Γ c,P remains convex in the standard linear geometry of probability measures while limiting the contribution of large transport displacements.
The paper develops the optimal-transport theory associated with the Huber ground cost. Despite the loss of strict convexity in its linear regime, the paper establishes:
-
Regularity and stability of Kantorovich potentials: Lemma 2.2 shows that dual potentials are c-Lipschitz. Theorem 2.5 establishes that T ρ c(µ1, ν) − T ρ c(µ2, ν) ≤ c · W1(µ1, µ2).
-
Uniqueness of potentials: Theorem 2.7 shows that under standard assumptions (µ ≪ Ld, µ(∂supp(µ)) = 0, supp(µ) connected), solutions of the dual problem are unique up to additive constants.
-
Existence of Monge solutions: Theorem 2.8 proves that for absolutely continuous source measures, there exists a Huber optimal transport map H c such that π H c = (I × H c)#µ.
-
Quantitative convergence toward quadratic optimal transport as c → ∞: Proposition 2.11 and Theorem 2.12 provide explicit bounds.
The paper proves:
-
Existence and convexity: Theorem 3.4 shows that m c(P) is non-empty, bounded, convex, and closed in W1.
-
Consistency under one-stage and two-stage sampling: Proposition 3.5 establishes almost sure convergence of empirical barycenters.
-
Characterization through averaged Kantorovich potentials: Theorem 3.7 characterizes minimizers through jointly measurable functions satisfying specific conditions.
The robustification parameter provides a genuine interpolation:
-
As c ↓ 0: Barycenters converge to W1-type Wasserstein medians (Theorem 3.6(i))
-
As c ↑ ∞: Barycenters converge to classical quadratic Wasserstein barycenters (Theorem 3.6(ii))
-
Finite-sample breakdown point: Theorem 3.11 shows that BP(m c(µ)) ∈ [⌈n/2⌉/n, (⌊n/2⌋+1)/n], essentially equal to 1/2.
-
Influence function in dimension one: Proposition 3.13 derives the pointwise influence function:
IF c(η; P)(u) = ψ c(Q η(u) − Q ν c(u)) / a c(u)
where a c(u) = P(Q µ(u) − Q ν c(u) < c).
-
Asymptotic distribution: Proposition 3.14 establishes pointwise asymptotic normality.
-
Robustness–efficiency tradeoff: The paper quantifies the tradeoff, showing that the asymptotic relative efficiency with respect to the Wasserstein mean follows the classical Huber calibration (k ≃ 1.345 gives approximately 95% Gaussian asymptotic efficiency).
The paper shows that displacement-wise Huberization can retain first-order information lost when a single Huber weight is assigned to the distribution as a whole. Proposition 3.16 demonstrates this with localized contamination:
IF in c(η M,α; P0)(u) = c·1(1−α,1)(u)
IF out c(η M,α; P0)(u) = (c/√α)·1(1−α,1)(u)
This difference is analogous to cellwise versus casewise robustness in classical robust statistics.
Theorem 2.1 (Duality): For µ, ν ∈ P1(Rd), there exists an optimal plan π*, strong duality holds, and dual solutions consist of ρ c-conjugate, ρ c-concave functions.
Lemma 2.2: Any proper ρ c-concave function is Lipschitz with constant c, and its domain is all of Rd.
Theorem 2.5 (Stability of the cost): T ρ c(µ1, ν) − T ρ c(µ2, ν) ≤ c · W1(µ1, µ2).
Theorem 2.7 (Uniqueness of potentials): Under mild assumptions on µ, Sol*2(ν, µ)/∼ is a singleton.
Theorem 2.8 (Existence of transport maps): For µ, ν ∈ P1(Rd) with µ ≪ Ld, there exists a Huber OT map H c pushing µ forward to ν.
Proposition 2.11: 0 ≤ ½W22(µ, ν) − T ρ c(µ, ν) ≤ ∫(∥x∥2 − c2/4)+ dµ(x) + ∫(∥y∥2 − c2/4)+ dν(y).
Theorem 2.12 (Stability of the OT map): If H0 is α-Lipschitz, then ∥H0 − H c∥ L2(µ) ≤ 2α[∫(∥x∥2 − c2/4)+ dµ(x) + ∫(∥y∥2 − c2/4)+ dν(y)] 1/2.
Proposition 3.1: For every µ, ν ∈ P1(Rd), T ρ c(ν, µ) − c∫∥x∥ dµ(x) ≤ c∫∥y∥ dν(y) + c2/2.
Theorem 3.4 (Existence): For any P ∈ P(P1(Rd)), m c(P) is non-empty, bounded, convex, and closed in W1.
Proposition 3.5 (Consistency): Under both one-stage and two-stage sampling, empirical barycenters converge almost surely to population barycenters.
Theorem 3.6 (Stability as c varies): As c ↓ 0, barycenters converge to W1-type Wasserstein medians; as c ↑ ∞, they converge to quadratic Wasserstein barycenters.
Theorem 3.7 (Characterization): ν is a Huber–Wasserstein barycenter if and only if there exists a jointly measurable function (µ, x) ↦ f ν,µ(x) such that:
(i) For P-a.e. µ, f ν,µ is a Huber potential for (ν, µ)
(ii) For ν-a.e. x, ∫f ν,µ(x) dP(µ) = 0
(iii) For every x ∈ Rd, ∫f ν,µ(x) dP(µ) ≥ 0
Theorem 3.11 (Breakdown point): BP(m c(µ)) ∈ [⌈n/2⌉/n, (⌊n/2⌋+1)/n].
Proposition 3.12: In dimension one, ν ∈ P1(R) is a Huber–Wasserstein barycenter if and only if:
∫ ψ c(Q ν(u) − Q µ(u)) dP(µ) = 0 for a.e. u ∈ (0, 1)
where ψ c(t):= max −c, min t, c.
Proposition 3.13 (Influence function): IF c(η; P)(u) = ψ c(Q η(u) − Q ν c(u)) / a c(u), with sup η∈P2(R) IF c(η; P)(u) ≤ c/a c(u).
Proposition 3.14 (Asymptotic normality): √n(Q ν̂ n,c(u) − Q ν c(u)) ⇒ N(0, (1/a c(u)2)∫ψ c2(Q µ(u) − Q ν c(u)) dP(µ)).
Proposition 3.15 (Coincidence under random translations): Under the random-translation model, inside- and outside-Huber procedures have identical influence functions and first-order asymptotic efficiencies.
Proposition 3.16 (Localized contamination): For M > c and √M·α > c, IF in c(η M,α; P0)(u) = c·1(1−α,1)(u) while IF out c(η M,α; P0)(u) = (c/√α)·1(1−α,1)(u).
Proposition 3.17 (Efficiency under localized shape variation): The asymptotic relative efficiency of the outside estimator with respect to the inside estimator for the unaffected lower-half location converges to 1 − ε as τ → ∞.
The first experiment uses nine handwritten digit images (six digit-1 images as target population, three digit-0 images as contaminants). The Wasserstein mean is visibly affected by the digit-0 pollutants, while the Wasserstein median remains closer to the vertical-stroke geometry shared by the majority digit-1 images. Varying c interpolates between these behaviors: large c gives mean-like behavior, small c gives median-like behavior.
The second experiment uses passenger-flow profiles for London Underground stations. Morning-peaked profiles form the clean target population, evening-peaked profiles are pollutants. The selected cutoff b̂c = 0.02 places the Huber center close to the median. Results show:
Estimator Validation W2 Test W2
Wasserstein Mean 0.049448 0.052597
Huber–Wasserstein mean, ĉ = 0.02 0.045468 0.051821
Wasserstein Median 0.045483 0.052161
The third experiment uses subgroup-level score distributions on [0,1] with clean and biased subgroups. The cutoff is selected via robust cross-validation using an 85%-trimmed W2 criterion. Results show:
Estimator Training-CV Trimmed W2 Mixed-Test Trimmed W2
Wasserstein Mean 0.155057 0.138132
Huber–Wasserstein mean, ĉ = 0.06 0.133948 0.109739
Wasserstein Median 0.135911 0.109907
The selected Huber center lies between the mean and median behaviors, discounting biased distributions while retaining more averaging than the median.
The paper includes a Huber–Sinkhorn algorithm (Algorithm 1) for computing regularized Huber–Wasserstein barycenters. The algorithm uses the Huber cost matrix C(c) with entries C(c) ij = ρ c(∥x i − x j∥) and kernel K(c,ε) = exp(−C(c)/ε). Theorem B.3 establishes existence and uniqueness of the regularized barycenter, and Theorem B.5 proves convergence of the Sinkhorn iterates.
The paper distinguishes its approach from several related methods:
-
ROBOT framework (Mukherjee et al., 2021): Uses hard-truncated ground cost C λ(x,y) = min C(x,y), 2λ, which completely saturates transportation costs above the threshold. The Huber cost instead clips the score (ρ′ c(r) = min r, c) while allowing the cost to grow linearly.
-
Hard-truncated ground metric (Cheng and Liu, 2026): Uses d λ(x,y) = min d(x,y), λ, which is a hard-rejection mechanism. The Huber cost is convex and continuously differentiable, retaining quadratic geometry for small displacements.
-
Metric-space Huber means (Lee and Jung, 2026): Apply the Huber loss outside the optimal transport problem, assigning a single scalar weight to each observation. The proposed approach clips individual transport displacements separately.
-
p-Wasserstein barycenters (Brizzi et al., 2025) and h-Wasserstein barycenters (Brizzi et al., 2026): The Huber cost falls outside the strictly convex framework because it is only strictly convex in its quadratic region and becomes affine beyond the cutoff c.
The paper also connects to research on the robustness of optimal transport maps themselves (Avella-Medina and González-Sanz, 2024; Paindaveine and Passeggeri, 2024; González-Sanz and Avella-Medina, 2026; González-Sanz et al., 2026), noting that the statistical object considered here is different: observations are probability measures, contamination occurs at the level of a population of distributions, and the object of study is their barycenter.
Improvements for AI systems
Based on the paper, here are specific improvements for AI systems:
Improvement: Replace standard FedAvg (which uses Wasserstein-2 barycenters) with Huber–Wasserstein barycenters for aggregating client model distributions.
What the improved system can do:
-
Automatically down-weight clients with corrupted or adversarial model updates (e.g., poisoned gradients) while preserving the geometric structure of legitimate updates
-
Maintain robustness against up to 50% malicious clients (breakdown point ≈ 1/2) without explicit outlier detection
-
Tune the Huber parameter c to balance between aggressive averaging (large c) and median-like robustness (small c) based on the trust level of the client population
Sources
- On the breakdown point of transport-based quantiles
- Wasserstein Spatial Depth
- Robust Wasserstein barycenter
- Banach-Saks Theorem for $L^1$ revisited
- Breakdown properties of optimal transport maps: general transportation costs
- The Influence Function of Transport-based Quantiles
- On the robustness of semi-discrete optimal transport
Related papers
- Doubly robust inference via calibration
- Bayesian Empirical Bayes: Simultaneous Inference from Probabilistic Symmetries
- Flexible Nonparametric Inference for Causal Effects under the Front-Door Model
- Deployment of AI-Assisted Interventions: Capacity Constraints and Noisy Compliance
- A Survey on Archetypal Analysis
- Dynamic Spatial Bayesian Machine Learning Model: Applications to Intergenerational Economic Mobility and Geographic Income Inequality in the United States