Statistical Properties of Robust Learning under Distributional Shifts

arXiv:2608.13133 · stat.ML, cs.LG · Submitted 2026-08-13 · Read on arXiv

Zhiyi Li, Xiaojie Mao, Yunbei Xu, Ruohan Zhan

National University of Singapore · Tsinghua University · University College London

stat.ML, cs.LG

Submitted: 2026-08-13

Updated: 2026-08-14

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 75/100

The gist: This paper studies the statistical properties of robust learning methods—specifically Distributionally Robust Optimization (DRO) and Robust Satisficing (RS)—when there are distributional shifts

Terminology

Summary

This paper studies the statistical properties of robust learning methods—specifically Distributionally Robust Optimization (DRO) and Robust Satisficing (RS)—when there are distributional shifts between the source environment (where training data is generated) and the target environment (where the model is deployed). The authors note that Distributional shifts arise when the target deployment environment differs from the source environment that generated the training data. While robust learning frameworks like DRO and RS aim to address this challenge, "their finite-sample guarantees under such shifts, and their systematic comparison, remain underexplored: existing analyses typically establish guarantees either in the source environment or for adversarial worst-case performance over an ambiguity set."

The paper's central contribution is studying generalization error in the target environment—the excess loss under the shifted target distribution. The authors state their contributions are threefold: "First, we derive finite-sample generalization error bounds in the shifted target environment for both DRO and RS. These bounds explicitly characterize the trade-off between reduced sensitivity to shift and the regularization penalty induced by each method's robustness hyperparameter, and they avoid the curse of dimensionality associated with Wasserstein empirical concentration. Second, when partial shift information such as shift magnitude or direction is available, we propose information-directed hyperparameter calibrations and compare the two methods given the same information. Under these calibrations, and in the partial-information regimes we study, DRO and RS exhibit complementary theoretical and empirical behavior. Finally, we apply the framework to a network lot-sizing problem, using it to interpret how robust policies respond to positive shifts in the demand distribution."

The paper defines the generalization error under distributional shifts as follows: "Given n i.i.d. samples zi ni=1 drawn from a source distribution PS, an algorithm learns an estimated decision x̂ ∈ X. Then x̂ is evaluated by its generalization error, defined as the excess loss under the target distribution: RPT(x̂):= EPT[f(x̂, z)] − JT," where JT is the minimal expected loss under the target distribution.

The paper considers three methods:

  1. ERM (baseline): min EP̂n[f(x, z)] where P̂n is the empirical distribution

  2. DRO: min max EP[f(x, z)] over a Wasserstein ball B(P̂n, r) of radius r around the empirical distribution

  3. RS: formulated as kτ = min k s.t. EP[f(x, z)] − τ ≤ k·dW(P, P̂n), ∀P ∈ P(Z), x ∈ X, k ≥ 0

The paper relies on Assumption 1 (Regularity): "(a) Bounded Z. The instance space Z is bounded: diam(Z) = supz,z′∈Z z − z′ < ∞. (b) Bounded functions. The loss function f(x, z) is lower semicontinuous in x ∈ X and is uniformly bounded, i.e., 0 ≤ f(x, z) ≤ M < ∞, ∀x ∈ X, z ∈ Z. (c) Lipschitz continuity of loss. The loss function is Lipschitz in z, uniformly over x ∈ X."

Proposition 1 provides the ERM generalization bound: RPT(x̂ERM) ≤ L · dW(PS, PT) + [JS − JT] + (24/√n)C(A) + 2M√(log(2/δ)/(2n)), ∀PT. This bound consists of three components: "The first, L · dW(PS, PT), is the sensitivity-to-shift term, which quantifies the discrepancy between the source and target distributions. The second, JS − JT, represents the difference between the minimal expected losses under the source and target distributions, which we refer to as the environment gap. The third is the statistical learning error, a residual term that decreases at the standard n−1/2 rate."

Theorem 1 establishes the DRO bound: RPT(x̂DRO) ≤ L · inf P∈B(PS,r) dW(PT, P) + Λr(PS, xS) + [JS − JT] + (48/√n)C(A) + (72L·diam(Z))/n + 2M√(log(4/δ)/(2n)), ∀PT.

The paper explains: "This upper bound consists of four components. The first is the sensitivity-to-shift: L · inf P∈B(PS,r) dW(PT, P), which measures the discrepancy between the target distribution and the closest distribution in the Wasserstein ball centered at the source distribution. Setting r = 0 recovers the shift term in the ERM bound (8), while r > 0 allows DRO to improve upon it. In particular, once r ≥ dW(PS, PT) so that the target distribution PT lies in the Wasserstein ball B(PS, r), this shift term becomes zero. The second is the regularization penalty: Λr(PS, xS), which is defined in (9) and does not appear in the ERM bound. This term reflects the additional cost introduced by robustness; it vanishes when r = 0, aligning DRO with ERM."

Remark 1 states: "Theorem 1 provides the first target-environment generalization error bound in the literature that explicitly characterizes the trade-off induced by the DRO radius r. The shift term L · inf P∈B(PS,r) dW(PT, P) decreases as r increases, because enlarging the Wasserstein ball makes it more likely to cover the target distribution PT... In contrast, the regularization term Λr(PS, xS) grows with r, since a larger Wasserstein ball amplifies the worst-case loss deviation."

Theorem 2 establishes the RS bound: RPT(x̂RS) ≤ kτϵ · dW(PS, PT) + ϵ + [JS − JT] + (48/√n)C(A) + (48L·diam(Z))/n + 2M√(log(2/δ)/(2n)), ∀PT.

The paper explains: "This upper bound consists of four components. The first is the sensitivity-to-shift kτϵ · dW(PS, PT), which quantifies the distributional discrepancy. Unlike ERM, the RS framework introduces a hyperparameter-dependent multiplicative factor kτϵ in the shift term. By Lemma 1, this factor kτϵ is no larger than the Lipschitz constant L, thereby improving on the ERM bound. The second is the regularization penalty, ϵ, which is the tolerance value specified in the RS model (6) and reflects the additional cost of robustness."

Remark 2 notes: "Theorem 2 highlights the trade-off in setting the RS hyperparameter ϵ. On one hand, enlarging ϵ (and hence τ) reduces the fragility parameter kτϵ, thereby decreasing the coefficient of the shift term. On the other hand, the satisficing regularization penalty grows with ϵ. Thus, RS achieves robustness by attenuating the sensitivity coefficient on the source-target distance, whereas DRO achieves robustness by reducing the effective distance from the target distribution to the ambiguity set."

Remark 3 states: "We improve upon Li et al. (2024) by avoiding the Wasserstein concentration rate that leads to the curse of dimensionality in their bound. In our result, the statistical error depends on the instance space through diam(Z)/√n instead of a direct concentration bound for dW(PS, P̂n)... the sample-size rate remains O(n−1/2), improving over the rate O(n−min 1/d,1/2) that arises from direct Wasserstein concentration."

The paper provides a Dirac example with loss f(x, z) = 1 − xz, X = [0, 1], PS = δ1. Proposition 2 shows sharpness of the ERM bound: RPT(x̂ERM) = [JS − JT] + L·dW(PS, PT). Table 1 summarizes the actual risks and bounds for each method, showing DRO trades residual shift for radius regularization, whereas RS trades source-side tolerance for a smaller shift sensitivity.

The paper compares DRO and RS through trade-off terms defined as: TODRO(r) = SenDRO(r) + RegDRO(r), TORS(τ) = SenRS(τ) + RegRS(τ), where the sensitivity and regularization components are defined in Table 2.

When the shift magnitude r = dW(PS, PT) is known but direction is unknown, the paper proposes calibrating DRO with radius r and RS with threshold τr = supP∈B(P̂n,r) EP[f(x̂ERM, z)].

Proposition 3 shows: 0 = SenDRO(r) ≤ SenRS(τr) = kτr·r, RegRS(τr) − ρn(δ) ≤ RegDRO(r) ≤ RegRS(τr) + ρn(δ). The paper concludes: "Combining these two components, the aggregate trade-off comparison in Proposition 3 shows that DRO is favored by the non-vanishing sensitivity gap kτr·r, up to the statistical error ρn(δ). This stems from DRO's ability to directly incorporate the shift magnitude into its ambiguity set and absorb the target distribution into its worst-case robustness framework."

When the shift direction is known through a distribution family Pt t≥0 but magnitude is unknown, the paper proposes calibrating DRO with rt = dW(PS, Pt) and RS with τt = EPt[f(x̂ERM, z)].

Proposition 4 shows: RegRS(τt) + sup P∈B(PS,rt) EP[f(xS, z)] − EPt[f(xS, z)] ≤ RegDRO(rt) + ρ̄n(δ), meaning the DRO regularization penalty is asymptotically larger than the RS regularization penalty by a nonnegative gap.

Proposition 5 provides sensitivity comparisons: (i) If dW(Pt,PS)/dW(PT,PS) ≤ 1 − kτt/L, then SenRS(τt) ≤ SenDRO(rt); (ii) If t ≥ tT, then 0 = SenDRO(rt) ≤ SenRS(τt).

The paper explains: "Part (i) corresponds to cases where the distribution shift magnitude is sufficiently under-specified... Combining it with Proposition 4, the resulting trade-off bound favors RS up to the statistical error ρ̄n(δ). In contrast, part (ii) of Proposition 5 favors DRO for the sensitivity-to-shift term, when the shift magnitude is well-specified or over-specified and thus the ambiguity set already covers the true target distribution."

The paper presents a two-product Gaussian newsvendor problem with loss f(x, z) = Σj=12 [hj(xj − zj)+ + bj(zj − xj)+], with source distribution PS = N(µS, Σ) and mean-shift targets PT = N(µT, Σ).

Scenario I results: "DRO has overall lower realized target loss than RS across the reported magnitudes, which is consistent with Proposition 3: when only the shift magnitude is available, DRO is favored as it uses the magnitude directly through its ambiguity radius."

Scenario II results: The paper considers high alignment (shift direction close to e2, where product 2 has larger underage cost) and low alignment (shift direction close to e1). "In the high-alignment regime... When the shift magnitude is largely under-specified, both robust methods can improve upon ERM in this regime, and RS can be better than DRO. When the magnitude is over-specified, DRO in turn becomes more favorable than RS."

The paper applies the framework to a network lot-sizing problem with N=10 stores, where the total system cost consists of three components: the initial-ordering cost, the transportation cost from transshipment, and the emergency ordering cost. The two-stage model is defined in equation (15).

Key findings: "In the under-specified regime (t < tT), robust methods generally improve upon ERM because the source-trained ERM decision tends to under-order for upward demand shifts. With large under-specification, RS attains a lower total cost than DRO. However, The performance changes in the over-specified regime (t > tT)... When the unit initial-ordering cost ci is small, DRO becomes highly conservative and performs substantially worse than both RS and ERM. As ci increases, this disadvantage diminishes, and DRO can eventually outperform RS."

The cost decomposition shows: "Relative to ERM, both DRO and RS typically incur higher initial-ordering costs but lower operational costs, indicating that robust methods substitute preventive upfront inventory for corrective second-stage transshipment and emergency ordering."

The paper also compares optimization-based correspondence (Wang et al. 2025) with shift-calibrated hyperparameter pairs, showing how the relative performance of DRO and RS changes with the initial-ordering cost c.

The paper concludes: "We derive generalization error bounds for both methods under distributional shifts, explicitly characterizing the trade-off between reduced sensitivity to shift and regularization cost as a function of their respective hyperparameters... Under such hyperparameter calibration, when the shift magnitude is known but the direction is unknown, the resulting trade-off bound favors DRO because the ambiguity radius can directly encode the known magnitude. When the shift direction is known but the magnitude is unknown, RS can be favored when the magnitude is sufficiently under-specified."

Improvements for AI systems

Improvement 1: Shift-Aware Hyperparameter Auto-Calibration

The AI system can automatically calibrate its robustness hyperparameters (DRO radius r or RS tolerance ϵ) based on available partial shift information. When shift magnitude is known but direction is unknown, it defaults to DRO with radius set to the known magnitude. When shift direction is known but magnitude is uncertain, it switches to RS with tolerance calibrated to the direction family, favoring RS when magnitude is under-specified and DRO when over-specified. This enables the system to select the optimal robust learning method and hyperparameters without manual tuning, improving deployment performance under distributional shifts.

Improvement 2: Shift-Sensitivity-Aware Model Selection

The AI system can compare candidate models (ERM, DRO, RS) using the derived trade-off bounds (sensitivity-to-shift + regularization penalty) rather than only validation loss on source data. Given an estimated shift magnitude and direction, it can compute the theoretical target-environment generalization bound for each method and select the one with the lowest predicted excess loss. This prevents over-reliance on ERM when shifts are large, and avoids over-conservative DRO when shifts are small or well-specified.

Improvement 3: Finite-Sample Guarantee-Aware Decision Making

The AI system can provide confidence intervals on its target-environment performance using the derived O(n-1/2) bounds that avoid the curse of dimensionality. For safety-critical applications (e.g., inventory management, medical decisions), it can report worst-case target loss with a specified confidence level, accounting for both statistical error and distributional shift. This enables risk-aware decision-making where the system can reject actions whose worst-case target loss exceeds a safety threshold.

Improvement 4: Robustness-Regularization Trade-off Visualization

The AI system can generate an interactive trade-off curve for any given task, showing how target generalization error varies with robustness hyperparameters (r for DRO, ϵ for RS). It decomposes the error into sensitivity-to-shift and regularization penalty components, allowing users to visually identify the optimal operating point. This helps practitioners understand whether their current robustness setting is over- or under-regularized relative to the actual shift magnitude.

Improvement 5: Shift-Information-Adaptive Learning Pipeline

The AI system can dynamically adjust its learning strategy based on the type of shift information available: (a) no shift information → use ERM with standard regularization; (b) shift magnitude known → use DRO with radius equal to magnitude; (c) shift direction known → use RS with tolerance calibrated to direction family, preferring RS when magnitude is under-specified; (d) both magnitude and direction known → use DRO with radius slightly above magnitude to ensure coverage. This creates a principled decision tree for robust learning that maximizes target performance given the available information.

Improved AI System Capabilities:

The enhanced AI system can:

  • Automatically select between ERM, DRO, and RS based on available shift information, with theoretical guarantees on target-environment performance

  • Provide finite-sample confidence bounds on deployment performance that explicitly account for distributional shift, avoiding the curse of dimensionality

  • Calibrate robustness hyperparameters (r, ϵ) to match the estimated shift magnitude/direction, balancing sensitivity reduction against regularization cost

  • Predict when DRO will outperform RS (known magnitude, unknown direction) and vice versa (known direction, under-specified magnitude), and switch accordingly

  • Decompose target loss into interpretable components (shift sensitivity, regularization penalty, environment gap, statistical error) for debugging and explanation

  • Apply these capabilities to supply chain problems (e.g., lot-sizing under demand shifts) where it can recommend robust ordering policies that trade off upfront inventory against corrective actions, with quantified worst-case cost guarantees

Sources

Related papers