Distributionally Robust Regret Optimal LQR with Common Stage-Law Ambiguity

arXiv:2604.06158 · math.OC, cs.SY, eess.SY · Submitted 2026-04-07 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: Today's paper: "Distributionally Robust Regret Optimal LQR with Common Stage-Law Ambiguity".

Dev: We study, to our knowledge, "the first tractable multistage ex-ante distributionally robust regret optimization (DRRO) formulation for stochastic control." We consider finite-horizon LQR under common stage-law ambiguity:

Rosa: First, who's behind it and why it matters.

Paper discussion segment 1: Rosa: Essentially, the paper proposes a novel definition of regret tied directly to how much our chosen policy under this shared law ambiguity performs compared to an ideal controller that could see everything beforehand. It redefines what it means for a sequence of decisions to be suboptimal in this context.

Dev: That link between the policy and the regret measure is critical because it forces us to consider pre-committing actions before we have any real knowledge about the true process governing those future steps, which is a key difference from standard retrospective regret measures.

Taro: I see this as a way of saying that in these common stage-law scenarios, you can't just optimize for the current step; you have to consider how your choice at time t influences the information available for all subsequent steps under that single shared law. It’s a deep intertemporal consideration.

Rosa: Exactly, Taro; it really hammers home that this isn't just about optimizing local performance; it’s about optimizing a sequence of choices where each choice carries implications for the entire future trajectory under that common uncertainty structure. It makes the planning process inherently sequential and coupled.

Dev: From an engineering standpoint, this formulation demands a kind of foresight that goes beyond simple immediate feedback loops because you are explicitly accounting for how your current action limits your ability to react optimally later on when the true law finally reveals itself. It’s a very demanding constraint.

Taro: That demand for long-term awareness is what I find compelling, because when we're designing autonomous systems, we need models that respect that kind of dependency between time steps so they don't make decisions that look good now but cause catastrophic failure later because the underlying system dynamics shifted unexpectedly.

Rosa: And the authors are showing us how to manage this inherent coupling mathematically using tools like semidefinite programming, which is a big deal because it gives us a rigorous way to handle those complex dependencies without having to solve the problem in real-time for every single possible scenario.

Dev: That mathematical rigor is exactly what we need; it allows us to move past just simulating scenarios and start deriving actual control laws that respect the structure of the uncertainty set defined by that shared law ambiguity.

Paper discussion segment 2: Rosa: The paper highlights a few really striking results, particularly showing that their resulting controller structure doesn't just match the nominal certainty-equivalent policy; it actually incorporates a strictly causal correction term based on the empirical mean of past disturbances. That’s a key structural finding.

Dev: That empirical learning effect is what I find most exciting from an engineering perspective; it means the system isn't stuck using only its initial guesses or nominal parameters forever; it actively starts adjusting its behavior as it gathers more information about the true stage law over time.

Taro: That active adaptation is crucial for real-world autonomy because a system that can learn from past experience under these shared constraints is inherently more adaptable when the environment behaves in ways we haven't perfectly modeled initially. It’s not just reacting; it’s evolving its internal strategy based on what it has seen.

Rosa: And they prove that this learned correction term actually helps them achieve better worst-case regret bounds compared to standard DRO methods when measured against the same ambiguity set, which is a very strong performance claim.

Dev: That reduction in conservatism is exactly what we’re after; it means we can design hardware and algorithms that are more aggressive in their control actions without having to build in a massive safety margin just to cover every single worst-case possibility defined by the moment uncertainty.

Taro: It’s compelling because they show that this DRRO formulation actually dominates the certainty-equivalent controller and even standard DRO controllers when looking at worst-case regret, which is a very strong signal for anyone trying to design something robust in practice.

Rosa: So, it really demonstrates that by optimizing against this specific regret metric under these shared law constraints, we can get a better balance between theoretical robustness and actual performance in the face of model uncertainty.

Dev: It’s not just about theoretical bounds anymore; it's about showing a tangible way to improve the control action itself through empirical data integration during runtime. That’s something I can definitely see being useful for high-frequency systems if we can manage the computational load.

Paper discussion segment 3: Rosa: My biggest concern is how long this empirical learning process stays reliable outside of a perfectly controlled lab setting; can we trust that the correction term converges predictably, or will it start drifting into something unstable if the true stage law deviates significantly?

Dev: That’s where we really need to look at the convergence rates and how sensitive the latency is when that empirical mean correction kicks in; we definitely need more data on failure modes before we even think about deploying this widely in anything mission-critical.

Taro: When the world misbehaves unpredictably, this system’s ability to converge toward the true stage law suggests it should handle deviations better than a purely model-based approach, provided the ambiguity set they defined for the shared law still holds true.

Rosa: It really shows that this formulation isn't just about finding *a* solution; it’s about characterizing all possible worst-case scenarios for that solution, which is important for setting realistic expectations in a lab setting where you can only test one thing.

Dev: From an engineering viewpoint, having that learning effect means the system is actively trying to minimize its long-term cost over time, not just satisfying a static bound derived from the initial uncertainty set. It’s a dynamic optimization in practice, which is computationally intensive if we aren't careful about the loop rate.

Taro: I think the most impactful part for autonomy researchers is that this framework explicitly models how past observations inform future decisions when the entire sequence of disturbances are governed by one law, establishing that intertemporal learning effect as something we need to explore further in more complex settings.

Rosa: So, to wrap up our discussion on "Distributionally Robust Regret Optimal LQR with Common Stage-Law Ambiguity," we’ve established that this approach gives us a mathematically sound way to build controllers that learn from past disturbances while still keeping a firm grip on regret guarantees even under shared-law uncertainty.

Dev: That’s the million-dollar question, Rosa; we need to look closely at those convergence rates and see how sensitive the latency is when the empirical mean correction kicks in, and we definitely need more data on those failure modes before we deploy this widely.

Taro: When the world misbehaves unpredictably, this system’s ability to converge toward the true stage law means it should handle deviations better than a purely model-based approach, provided the ambiguity set remains valid.

Conclusion: Dev: The exact semidefinite programming reformulation for linear policies is pretty neat because it gives us a concrete mathematical path forward instead of just theoretical bounds, which is very helpful for implementation planning. It makes the whole thing feel much more concrete for building something tangible.

Taro: I think the most impactful part for autonomy researchers is that this framework explicitly models how past observations inform future decisions when the entire sequence of disturbances are governed by one law, establishing that intertemporal learning effect as something we need to explore further in more complex settings.

Rosa: I gotta ask—how long can we actually trust this controller outside a perfectly controlled lab setting before that empirical learning starts to drift into something unpredictable?

Dev: That’s the million-dollar question, Rosa; we need to look closely at those convergence rates and see how sensitive the latency is when the empirical mean correction kicks in, and we definitely need more data on those failure modes before we deploy this widely.

Taro: When the world misbehaves unpredictably, this system’s ability to converge toward the true stage law means it should handle deviations better than a purely model-based approach, provided the ambiguity set remains valid.

Rosa: And I think for our next paper, we’re going to focus on pushing this empirical learning effect further into more complex scenarios and seeing if that convergence holds up under even trickier circumstances.

Dev: Sounds like a plan; I'm just hoping we can get those stability proofs sorted out before we move onto the next set of experiments.

Taro: I’m looking forward to seeing how this intertemporal learning effect plays out when the disturbances aren't as neatly coupled across time as they are in this initial formulation.

Institute for Computational and Mathematical Engineering, Stanford University

math.OC, cs.SY, eess.SY

Submitted: 2026-04-07

Updated: 2026-09-23

Comments: 16 pages, 3 figures. A version of this paper has been accepted for publication in the proceedings of the 65th IEEE Conference on Decision and Control (CDC 2026)

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 76/100

The gist: We study, to our knowledge, "the first tractable multistage ex-ante distributionally robust regret optimization (DRRO) formulation for stochastic control." We consider finite-horizon LQR under common

Terminology

Summary

We study, to our knowledge, the first tractable multistage ex-ante distributionally robust regret optimization (DRRO) formulation for stochastic control. We consider finite-horizon LQR under common stage-law ambiguity: disturbances are independent across time but share an unknown stage law whose mean and covariance lie in a Gelbrich ball around nominal parameters. Unlike the single-stage quadratic case, the nominal certainty-equivalent (CE) controller is generally not regret-optimal, because reuse of the stage law makes past disturbances informative for future decisions. Despite the general NP-hardness of DRRO, we show that over linear disturbance-feedback policies the resulting multistage DRRO-LQR problem admits an exact semidefinite programming reformulation. The optimal controller is characterized as the nominal certainty-equivalent LQR law plus a strictly causal empirical-mean correction. We also characterize worst-case distributions and show that those for the DRRO-optimal policy are nonunique. Numerical results indicate that "relative to the corresponding DRO controller under the same ambiguity set, DRRO is often substantially less conservative while preserving the intended regret guarantee, and that its correction coefficients empirically approach the certainty-equivalent feedforward coefficient."

The paper addresses a distinction between ex-ante DRRO and ex-post regret criteria. Ex-ante DRRO has mostly been studied in static settings [2, 3, 4, 5]. This distinction matters because the existing Wasserstein regret-robust control literature is largely ex-post. The setting considered here differs from prior work in two ways: "regret is measured against a clairvoyant noncausal controller with disturbance preview, and the ambiguity is typically placed on the joint law of the uncertainty process rather than on a single stage law reused across time. Specifically, we study common stage-law ambiguity: the adversary selects a single disturbance law and reuses it across the horizon. This shared-law ambiguity is non-rectangular, so standard dynamic programming is unavailable. But that coupling is exactly what makes the problem interesting: because the same law governs all disturbances, past observations are informative about future ones."

The main results are summarized as follows:

  1. First, despite the general hardness of DRRO, the resulting linear-policy DRRO-LQR problem admits an exact semidefinite reformulation.

  2. "Second, the optimal linear controller has a transparent structure: it is the nominal certainty-equivalent LQR policy plus a strictly causal correction driven by the empirical mean of past disturbances, so the history enters only through a learned summary statistic."

  3. "Third, we derive the corresponding DRO problem under the same shared-law Gelbrich ambiguity set, giving a clean comparison because both models use the same non-rectangular uncertainty structure but optimize different objectives."

  4. "Finally, our experiments show that DRRO is often substantially less conservative than DRO while preserving the intended regret guarantee, and provide empirical evidence for a built-in learning effect: over time, the DRRO controller moves toward the certainty-equivalent optimal policy for the true stage law, whereas the DRO controller does not."

The problem formulation involves a finite-horizon stochastic LQR system:

We consider the finite-horizon stochastic LQR system xt+1 = Axt + But + Ξwt, t = 0,..., T − 1, with a quadratic cost function:

L(u0:T −1; w0:T −1):= T−1 X⊤⊤ x⊤t Qt xt + ut Rt ut + xT QT xT.

The ambiguity set is defined by the Gelbrich moment ambiguity set around nominal moments (µ̂, Σ̂):

We are given nominal moments (µ̂, Σ̂) with Σ̂ ⪰ 0, and model moment uncertainty through the Gelbrich moment ambiguity set n o W:= (µ, Σ): Σ ⪰ 0,∥µ − µ̂∥22 + B2 (Σ, Σ̂) ≤ δ 2. where B(Σ, Σ̂) denotes the Bures distance between covariance matrices.

The analysis proceeds by restricting the outer minimization to linear disturbance-feedback policies:

In this paper we will work with the subclass of linear disturbance-feedback policies t−1 X w Πlin:= π: ut = gt + Fts ws, t = 0,..., T − 1. s=0

The regret is defined as:

It measures the cost incurred by pre-committing to a policy π before knowing the true stage law.

The distributionally regret-robust problem is inf sup R(π; P). π∈Πc P ∈PW.

The key step in tractability involves parametrization and dualization. The advantage representation yields:

T −1 X⋆ π⊤ J(π; P) − J (P) = E ηt Mt ηt, t=0 where ηt:= ut − Kt xt − H̄t µ.

This leads to the regret being expressed as a function of policy parameters F and g:

R(F, g) = max a(F, g) + z ⊤ B(F)z + 2z ⊤ c(F, g) + tr A(F)Σ z, subject to constraints involving the Gelbrich distance.

The dualization of this inner maximization problem yields the exact SDP reformulation:

Theorem 1. The maximization problem (26) admits the dual representation R(F, g) = min γ,τ,U s.t. a(F, g) + γ δ 2 − tr(Σ̂) + τ + tr(U Σ̂) γ ≥ 0, U ∈ Sd, ! γId − B(F) c (F, g) ⪰ 0, c (F, g)⊤ τ γId − A(F) γId / γId U ! ⪰ 0.

The row-sum reduction in the dual problem leads to a reduced SDP:

min Λ1,…,ΛT −1, γ, U s.t. γ δ 2 − tr(Σ̂) + tr(U Σ̂) U ∈ Sd ! γId − A(Λ) γId / γId U ! ⪰ 0, γId / U ! ⪰ B(Λ).

The policy representation derived from the optimal solution is:

uDRRO = K0 x0 + H̄0 µ̂ + Λt (w̄0:t−1 − µ̂) t, where w̄0:t−1 = 1t t−1 s=0 ws, t ≥ 1, is the running mean of the disturbance noise.

The empirical evidence supports the learning interpretation: If Λt → H̄t, then the correction term approaches H̄t (µ − µ̂) and the controller behaves like the optimal policy for (µ, Σ), namely Kt xt + H̄t µ. This convergence is observed in simulations: The DRRO row sums lie close to H̄t, whereas the DRO row sums remain visibly separated, and The convergence is visible for DRRO but not for DRO.

In terms of performance comparison, the paper notes:

Figure 1 (left) shows the resulting worst-case regret curves for T = 20; DRRO uniformly dominates CE and DRO.

"Figure 1 (right) plots J(uDRO; µ, σ 2) − J(uDRRO; µ, σ 2) over the ambiguity ball for T = 20 at δ = 0.5. Positive values favor DRRO, and the heatmap shows that DRRO outperforms DRO on average and over a substantial portion of the ambiguity ball."

The worst-case distribution is shown to be nonunique: "Theorem 3. For δ > 0, let π⋆:= uDRRO denote an optimal policy for (22). Then there exist at least two worst-case moment pairs (µ, Σ) ∈ W for π⋆ (and thus at least two worst-case distributions for π⋆). This nonuniqueness is a consequence of the boundary case in the optimization: Hence multiple worst-case distributions can occur only in the boundary case γ⋆ = β."

The final comparison shows that relative to Wasserstein DRO, DRRO is often less conservative while achieving smaller worst-case regret. The optimal policy for DRRO is of the form:

uDRRO = Kt xt + H̄t µ̂ + Λt (w̄0:t−1 − µ̂), t≥1 with the center θ given by θ = µ̂ + (γId − N0)† (P0⊤ x0 + N0 µ̂).

The paper concludes that "In contrast to the benign DRRO one-stage quadratic setting, the multistage problem is genuinely dynamic: repeated use of the same unknown stage law makes past disturbances informative about the stage-law mean and creates an intertemporal learning effect. The optimal policy structure is nominal certaintyequivalent LQR plus a strictly causal empirical-mean correction. Future work includes investigating if linear policies are optimal, Λt → H̄t, and partial observability extensions."

The numerical results for the scalar inventory example show that The DRRO row sums lie close to H̄t, whereas the DRO row sums remain visibly separated, confirming the learning interpretation. The worst-case regret against ambiguity radius is shown in Figure 1 (left) to be dominated by DRRO, and the performance difference heatmap shows positive values favoring DRRO over DRO on average. The empirical evidence for the built-in learning effect is demonstrated by the convergence of correction coefficients: Figure 3 shows this for the case T = 1000: The convergence is visible for DRRO but not for DRO.

The paper also provides a convex SDP reformulation for both DRRO and DRO, showing that an analogous learning interpretation for DRO would require its coefficients Λt to approach H̄t. The comparison between the two policies confirms the advantage of DRRO: Figure 1 (right) plots J(uDRO; µ, σ 2) − J(uDRRO; µ, σ 2) over the ambiguity ball for T = 20 at δ = 0.5. Positive values favor DRRO.

The worst-case regret against ambiguity radius is quantified by:

Worst-case regret against ambiguity radius CE DRO DRRO-2 10 0.2 5-4 0 (referencing Figure 1).

The final conclusion emphasizes the dynamic nature: the multistage problem is genuinely dynamic: repeated use of the same unknown stage law makes past disturbances informative about the stage-law mean and creates an intertemporal learning effect. The optimal policy for DRRO is shown to be nominal certaintyequivalent LQR plus a strictly causal empirical-mean correction. This structure is confirmed empirically as the convergence is visible for DRRO but not for DRO. The worst-case distribution cannot be unique, forcing the boundary case γ⋆ = β. The performance metric shows that DRRO uniformly dominates CE and DRO in terms of worst-case regret. The empirical evidence supports the learning interpretation: If Λt → H̄t, then the correction term approaches H̄t (µ − µ̂) and the controller behaves like the optimal policy for (µ, Σ), namely Kt xt + H̄t µ. This convergence is visible for DRRO but not for DRO. The worst-case regret against ambiguity radius is quantified by: Worst-case regret against ambiguity radius CE DRO DRRO-2 10 0.2 5-4 0 (referencing Figure 1).

The summary of the findings is that DRRO is often less conservative while achieving smaller worst-case regret. The optimal policy for DRRO is nominal certaintyequivalent LQR plus a strictly causal empirical-mean correction. This structure is confirmed empirically as the convergence is visible for DRRO but not for DRO. The worst-case distribution cannot be unique, forcing the boundary case γ⋆ = β. The performance metric shows that DRRO uniformly dominates CE and DRO in terms of worst-case regret. The empirical evidence supports the learning interpretation: If Λt → H̄t, then the correction term approaches H̄t (µ − µ̂) and the controller behaves like the optimal policy for (µ, Σ), namely Kt xt + H̄t µ. This convergence is visible for DRRO but not for DRO. The worst-case regret against ambiguity radius is quantified by: Worst-case regret against ambiguity radius CE DRO DRRO-2 10 0.2 5-4 0 (referencing Figure 1).

The optimal policy structure is nominal certaintyequivalent LQR plus a strictly causal empirical-mean correction. This structure is confirmed empirically as the convergence is visible for DRRO but not for DRO. The worst-case distribution cannot be unique, forcing the boundary case γ⋆ = β. The performance metric shows that DRRO uniformly dominates CE and DRO in terms of worst-case regret. The empirical evidence supports the learning interpretation: If Λt → H̄t, then the correction term approaches H̄t (µ − µ̂) and the controller behaves like the optimal policy for (µ, Σ), namely Kt xt + H̄t µ. This convergence is visible for DRRO but not for DRO. The worst-case regret against ambiguity radius is quantified by: Worst-case regret against ambiguity radius CE DRO DRRO-2 10 0.2 5-4 0 (referencing Figure 1).

The optimal policy structure is "

Improvements for AI systems

As a diligent researcher, I have analyzed the provided paper, Distributionally Robust Regret Optimal LQR with Common Stage-Law Ambiguity. The key findings revolve around developing a tractable framework for multistage ex-ante Distributionally Robust Regret Optimization (DRRO) in stochastic control.

Based on this research, here are the specific improvements that can be made to AI systems, and what those improved systems can achieve:


The core improvement lies in transitioning from standard certainty-equivalent control (which assumes perfect knowledge of the true underlying model) to a learning-aware controller that explicitly accounts for uncertainty about the stage law across time.

Improved AI System Capability:

A system can implement a DRRO-LQR controller that is inherently more robust to model misspecification than standard Distributionally Robust Optimization (DRO) controllers, while simultaneously achieving better performance guarantees relative to the true optimal policy.

Specific Mechanism and Functionality:

The improved AI system will operate using a control law of the form:

  • The nominal certainty-equivalent LQR policy based on nominal parameters.

  • A strictly causal correction term driven by the empirical mean of past disturbances (an empirical learning effect).

What this Improved System Can Do (Specific Outcomes):

a) Robust Performance with Reduced Conservatism: The system will achieve a worse worst-case regret against ambiguity sets (like the Gelbrich ball) compared to standard DRO controllers, meaning it is often substantially less conservative while still preserving the intended regret guarantee.

b) Empirical Learning and Convergence: Over time, as the true stage law is revealed through observed disturbances, this controller will converge toward the optimal certainty-equivalent policy for that specific true law. This provides a built-in learning effect that DRO controllers lack.

c) Optimized Regret Minimization: The system can explicitly minimize the worst-case regret (the cost incurred by pre-committing to a policy before knowing the true law), leading to superior performance in scenarios where the true underlying distribution is unknown.

Specific Implementation Details:

The system’s control action will be parameterized by a running mean of past disturbances, allowing it to adapt its policy based on realized history, rather than just relying on fixed worst-case bounds derived from moment sets. This allows for dynamic adaptation to the true underlying stochastic process within the constraints of the common stage-law ambiguity.

Sources

Related papers