Inverse Linear Quadratic Gaussian Games: Constrained Setting and Transferability

arXiv:2609.29321 · eess.SY, cs.SY · Submitted 2026-09-24 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: Today's paper: "Inverse Linear Quadratic Gaussian Games".

Dev: This work addresses finite-horizon inverse linear quadratic Gaussian (LQG) games in a constrained setting and explores transferability in an unconstrained setting.

Rosa: First, who's behind it and why it matters.

Title and authors: Dev: Moving on to the paper's summary, it really boils down to two main contributions: first, characterizing the set of cost parameters and optimal dual values that generate a generalized Nash equilibrium in a constrained setting, and second, proposing an algorithm to compute those very parameters.

Rosa: Exactly. Beyond just finding the parameters, they also tackle transferability in an unconstrained setting by bounding how much the cost value changes between policies derived from identified versus expert cost parameters when applied to different dynamics. That’s a substantial extension of their work.

Taro: The core idea seems to be that for finite-horizon inverse LQG games, you can pinpoint the exact cost structure and dual values associated with an observed equilibrium, which then allows us to use that knowledge for control or prediction on systems close to the original expert system.

Dev: They showed this works in constrained settings by identifying parameters that reproduce the generalized Nash equilibrium policies and trajectories in those scenarios using numerical simulations and real-robot experiments. That suggests the method is viable for systems where constraints are a factor, not just idealized environments.

Rosa: And they also demonstrated that when the setting is unconstrained, they can use these identified cost parameters to control sufficiently close dynamics, showing that the identified parameters are transferable under certain conditions.

Taro: It’s important to note what they explicitly state about their limitations; one of the limitations mentioned in their analysis is that the constant nu in Theorem one depends exponentially on the square of the horizon T squared, which gives us a specific scaling factor we have to account for.

Dev: That exponential dependence on T squared is something I need to keep an eye on when we look at loop rates and latency; that suggests that for very long horizons, the error bounds might become quite large unless our dynamics are extremely close.

Rosa: So, while they’ve established a solid framework for identification and transferability, we still have to be mindful of how those exponential factors scale with time in real-world operational deployments.

Taro: What I find most interesting is the link between the characterization using KKT conditions and the subsequent computation algorithm; it provides a rigorous pathway from observation to actionable parameters.

Dev: That algorithm involves solving a quadratic program for the cost parameters and then a linear complementarity problem for the dual variable, which means we’re dealing with two distinct mathematical optimization problems stacked together. I need to ensure our computational pipeline can handle that kind of complexity efficiently.

Rosa: So, they’ve provided both the theoretical foundation for what to look for and a practical roadmap on how to actually find it, which is valuable because it moves us from just theory to implementation.

Taro: This paper helps in understanding how complex multi-agent interactions can be simplified by identifying the underlying cost structure, even when we are operating under constraints.

The paper's summary: Rosa: Looking at the improvements suggested by this paper, it seems to focus on moving from just finding a solution to creating an algorithm that systematically computes the necessary parameters for any given generalized Nash equilibrium.

Dev: I agree. They propose an algorithm specifically designed for this purpose, which means we aren't just looking at a static set of equations; we are looking at a procedure that can actually generate those parameters from observations.

Taro: The improvement is moving toward an automated process where the system can ingest observed equilibrium policies or trajectories and output the cost parameters and dual values directly. That automates the identification process significantly.

Rosa: That’s what I mean; it shifts the focus from manual parameter tuning to a systematic, algorithmic approach for recovering those parameters in constrained LQG games that generate a specific Nash equilibrium.

Dev: This is exciting because it means we can automate the identification of these parameters, which is crucial if we need to apply this to many different systems or scenarios. The paper also extends this idea into unconstrained settings via transferability bounds.

Rosa: So, the improvements suggest that the paper’s real value isn't just in proving existence but in providing a robust and computationally tractable way to actually implement the identification process reliably across different dynamic conditions.

Taro: And I think the focus on transferability is key because it gives us confidence that if we identify parameters, we can use them even when the system dynamics deviate from the expert model by some margin.

Dev: That robustness is exactly what an engineer needs to hear; if our identified parameters are guaranteed to work on slightly perturbed dynamics, then our control design isn't overly sensitive to modeling errors.

Rosa: So, in essence, they’ve refined the method so that we have a systematic way to recover the cost structure and dual values, and then we have a way to verify its applicability outside of the exact training conditions.

The paper's improvements: Dev: Wrapping up this discussion on "Inverse Linear Quadratic Gaussian Games: Constrained Setting and Transferability," it seems the main implication is that we can recover the cost parameters governing multi-agent interactions from observed equilibrium policies or finite demonstrations.

Rosa: That’s a big statement—that the performance degradation when using these identified cost parameters scales linearly with deviations in dynamics and cost parameters, which is a practical metric for assessing reliability in real-world deployment.

Taro: I think the overall impact is that we get a tool to predict how much performance will degrade based on those deviations, bounded by O(epsilon T six N F squared T xi seven N xi K + E trace).

Dev: That bound, especially with the O(epsilon T six (N xi) two) scaling mentioned in Theorem one tells me that performance degradation is manageable as long as those errors epsilon stay small relative to the horizon T.

Rosa: It sounds like we can design and train policies for similar systems using these identified cost parameters, knowing that the performance will degrade predictably based on how far the actual dynamics stray from what was used during identification.

Taro: So, in a constrained setting, this paper provides a method to recover the underlying cost structure and dual values through tractable optimization problems, and in an unconstrained setting it offers transferability bounds for those parameters.

Dev: It’s definitely a solid piece of work that gives us concrete mathematical tools to bridge the gap between observed behavior and system design.

Rosa: I think this paper gives us a lot to consider as we move toward building more intelligent, adaptive systems that can handle uncertainty in complex multi-agent environments.

Conclusion: Taro: To conclude our discussion on "Inverse Linear Quadratic Gaussian Games: Constrained Setting and Transferability," the most significant takeaway is that the identified cost parameters can be used to predict or design policies for similar systems, with performance degrading linearly with the deviation of the dynamics and cost parameters.

Dev: That linear degradation bound means we have a quantifiable measure of how much our control loop will slip when things drift away from our identified model.

Rosa: So, we can confidently design and train policies that perform well even if the physical dynamics aren't perfectly matched to what was used for identification, provided those deviations are within the bounds established by the work.

Taro: This method provides a way to recover the cost structure and dual values from observed equilibrium policies or finite demonstrations, which is valuable for understanding complex multi-agent interactions under constraints.

Dev: It gives us concrete mathematical tools to bridge that gap between observation and system design, which is definitely something we can build on for more reliable control loops.

Rosa: I think this paper offers a lot to consider as we move toward building more intelligent, adaptive systems that can handle uncertainty in complex multi-agent environments.

SYCAMORE Lab, École Polytechnique Fédérale de Lausanne (EPFL)

eess.SY, cs.SY

Submitted: 2026-09-24

Updated: 2026-09-24

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 68/100

The gist: This work addresses finite-horizon inverse linear quadratic Gaussian (LQG) games in a constrained setting and explores transferability in an unconstrained setting.

Key concepts

Inverse LQG Games
These are games where the goal is to determine the underlying cost structure and dual values from observed equilibrium policies. The paper focuses on finite-horizon versions in constrained settings.
Transferability
This concept explores how much the cost value changes between policies derived from identified parameters and those applied to different system dynamics. It helps ensure that identified parameters are reliable even when system dynamics deviate slightly from the original model.
Cost Parameters and Dual Values
These are the specific mathematical structures that characterize a generalized Nash equilibrium in constrained LQG games. The paper provides a method to compute these parameters using optimization problems like quadratic programs and linear complementarity problems.
Linear Degradation Bound
This bound quantifies the performance degradation when using identified cost parameters on slightly perturbed dynamics. It shows that performance loss scales linearly with deviations in dynamics and cost parameters, providing a practical measure of reliability.

Terminology

Summary

This work addresses finite-horizon inverse linear quadratic Gaussian (LQG) games in a constrained setting and explores transferability in an unconstrained setting.

Contributions:

  1. For a finite-horizon constrained LQG game, the authors characterize the set of cost parameters and the optimal dual values corresponding to a given generalized Nash equilibrium and propose an algorithm to compute them.

  2. For inverse LQG games in an unconstrained setting, they derive a transferability bound on the cost value when the identified cost parameters with errors are applied to perturbed multi-agent dynamics.

Key Concepts and Characterization:

(The paper characterizes the generalized Nash equilibrium under constraints by considering optimality conditions such as Karush–Kuhn–Tucker (KKT) conditions [13], [14].)

(The characterization of the cost parameters and optimal dual value associated with the observed generalized Nash equilibrium is given in Proposition 1.)

Characterization Results (Proposition 1):

(i) For all players i ∈ [N] and t ∈ [T], the quadratic cost parameters and auxiliary linear term θ̄ti are characterized by:

"Mti θ̄ti = 0nu (nx +1)×1,

Qit ⪰ 0,

Rt−1 ≻ 0,

St−1

(9b) iSt−1

0nu ×n2x

0nu nx ×nx i⊤ Bt−1i∗⊤ −Kt−1 ⊗ I nu i⊤ −Et−1 ⊗ Inu i¯ it

St-1 ∆.i⊤ Bt-1omegāit"

(ii) The optimal dual variable λ∗ satisfies:

"˜ = 0,

λ∗ ⊤ (C̃λ∗ + d)∗

λ ≥ 0, and

C̃λ + d˜ ≤ 0."

Computation Algorithm:

(The authors propose an algorithm to compute the cost parameters and dual values based on Proposition 1.)

(The computation involves solving a quadratic program for the cost parameters and a linear complementarity problem for the optimal dual variable.)

Algorithm 1 Cost identification in LQG games (See Algorithm 1)

Transferability Analysis:

The notion of transferability [2] bounds the cost value perturbation between two policies: one induced by identified cost parameters, and the other by expert parameters.

(In an unconstrained setting, they derive a transferability bound on the cost value when identified cost parameters with errors are applied to perturbed multi-agent dynamics.)

Theorem 1. The identified cost parameters with errors θbid are ν-transferable to P (Definition 2), where

Empirical Validation:

The results are validated through numerical simulations and real-robot experiments:

(In a constrained setting, the algorithm is shown to identify cost parameters that recover the generalized Nash equilibrium policies in constrained LQG games.)

We show that our algorithm can identify cost parameters that recover the generalized Nash equilibrium policies in constrained LQG games.

(In an unconstrained setting, they show with a traffic simulation and real-robot experiments that the identified cost parameters can be used to control sufficiently close dynamics with performance degrading linearly with the deviations in the dynamics and the identified cost parameters.)

We also show that, under an unconstrained setting, the identified cost parameters are transferable to sufficiently close dynamics.

Theorem 1 (Transferability Bound):

"Theorem 1. The identified cost parameters with errors θbid are ν-transferable to P (Definition 2), where

P:= (Â, B̂)

max t∈[T −−], i∈[N] i E,i

E,i

max t∈[T −−], i in [N], E,i

Ât - AEt, B̂t - Bt ≤ ϵ"

(The cost value perturbation scales linearly with the errors of the identified cost parameters and the dynamics ϵ. However, the constant ν depends exponentially on the square of the horizon T2.)

"Theorem 1 shows that the cost value perturbation scales linearly with the errors of the identified cost parameters and the dynamics ϵ. However, the constant ν depends exponentially on the square of the horizon T2, inherited from the policy perturbation bound in Proposition 2."

Sample Complexity Result:

"Corollary 1. Suppose the Nash equilibrium policy used for cost identification in Algorithm 1 is estimated from n = O (ϵ−2 log N T /δ) demonstration samples. Then, with probability at least 1 − δ, the identified cost parameters are ν-transferable to P as per Theorem 1."

Conclusion:

The identified cost parameters can be used to predict or design policies for similar systems, with the performance degrading linearly with the deviation of the dynamics and cost parameters. The results indicate that the cost parameters governing multi-agent interactions can be recovered from observed equilibrium policy or finite demonstrations. The performance degradation is bounded by:

N ξ ∆K̄ + ∆ᾱ Jbi − Ji ≤ O T 6 N F̄t2T ξ 7 N ξ ∆K + Etrace.

This is bounded by O ϵ T 6 (N ξ 2)O(T), as stated in Theorem 1.

Index Terms:

Game theory; Identification; Stochastic control.

Notation Summary:

(A Gaussian distribution with mean µ and covariance matrix Σ is denoted as N (µ, Σ). The weighted L2 norm of a vector x with √a P ⪰ 0 is∥x∥P = x⊤ P x. Let In denote the identity matrix in Rn×n. We use diag(A1,..., Am) to denote the block-diagonal matrix with diagonal blocks A1,..., Am. The vectorization of A ∈ Rm×n is denoted by vec(A) ∈ Rmn. The Kronecker product of A ∈ Rm×n and B ∈ Rp×q is denoted by A ⊗ B ∈ Rmp×nq. Let f (·), g(·) be real-valued functions, we write f (ϵ) = O(g(ϵ)) if there exist constants c > 0 and ϵ0 such that f (ϵ) ≤ c g(ϵ) for all ϵ ≥ ϵ0.)

Appendix A Summary:

(Appendix A provides a proof of Lemma 1, which bounds the policy perturbations under cost parameter errors and dynamics errors. It involves complex recursions on Riccati parameters Pti and ζti, leading to bounds like:)

∥∆Kt ∥ = O N ξ 5 ϵ + N ξ 3 ∆Pt+1

∥∆αt ∥ = O N ξ 4 ϵ + N ξ 2 ∆Pt+1 + ξ squared ∆ζt+1

(Appendix B provides a proof of Lemma 2, which bounds the policy perturbation under fixed cost parameters and dynamic errors. It involves propagating Riccati perturbations backward and yields bounds such as:)

∥Kt′ − KtE ∥ = O N √ (T −t+2) 2

αt′ - αtE = O N ξ ϵ

(Appendix C derives the transferability cost bound in Theorem 1 by bounding the state trajectory perturbations and covariance trace terms, resulting in a final bound of:)

Jbi − J i ≤ O T 6 N T +5T + 2 ξ 2T +12T +16 ϵ.

Appendix D Summary:

(Appendix D defines the matrices used in the analysis, including:

"Φt ∈ RN nu ×N nu,

Ψt ∈ RN nx ×N nu,

x̂⊤ Qi x̂ − x⊤ Qi x ≤ Qi x̂ + x ∆x"

(It also defines the stacked matrices like F, Z, G, and Qit.)

Final Summary Statement:

"Given a generalized Nash equilibrium policy in a constrained setting, the cost parameters and optimal dual values can be identified through tractable quadratic programs. In an unconstrained setting, under small errors in the identified cost parameters and in a sufficiently close neighbourhood of the expert dynamics, the Nash policies and cost values remain close to those under the expert cost parameters. Through numerical simulations, we showed that the proposed method recovers the trajectories corresponding to the generalized Nash equilibrium in a constrained setting. Real-robot experiments further demonstrated the transferability of the identified cost parameters to sufficiently close dynamics in the unconstrained setting. The performance degrades linearly with deviations in dynamics and cost parameters. (Page 9)

Improvements for AI systems

Here are specific improvements to AI systems based on the provided research:

  • Identify unknown cost functions in multi-agent, finite-horizon Inverse Linear Quadratic Gaussian (LQG) games from observed policies or trajectories.

  • Compute the parameters (cost matrices, linear terms) and optimal dual values that generate a specific generalized Nash equilibrium policy under constrained settings.

  • Perform transferability analysis to determine how well identified cost parameters perform when applied to slightly perturbed dynamics (e.g., different vehicle models or environmental conditions).

This improved AI system can achieve the following:

  • Identify the underlying objective functions (cost metrics) that govern the behavior of multi-agent systems from recorded interactions, such as autonomous vehicles in a warehouse or coordinated robot fleets.

  • Reconstruct desired policies and trajectories for complex, constrained dynamic games (e.g., those requiring collision avoidance) even when the exact cost function is unknown.

  • Design and train control policies for similar systems using learned cost parameters that are guaranteed to perform close to the expert's performance, even if the physical dynamics change slightly (e.g., adapting a learned control law from one robot model to another).

  • Predict how performance will degrade (linearly) when the identified cost parameters are used on dynamics that deviate from those used during identification.

Abstract

This work addresses finite-horizon inverse linear quadratic Gaussian games. In a constrained setting, we characterize the set of cost parameters and optimal dual values that generate a given generalized Nash equilibrium, and we propose an algorithm to compute these parameters. In an unconstrained setting, we address transferability, namely, we bound the cost value perturbation between two policies: one induced by the identified cost parameters, the other by the expert parameters, under a set of different dynamics. This cost value perturbation scales linearly with the deviations in the dynamics and the identified cost parameters. Through numerical simulations, we show that, in a constrained setting, our algorithm identifies the cost parameters and dual values that can reproduce the policy and trajectories corresponding to the observed generalized Nash equilibrium. In an unconstrained setting, we show with a traffic simulation and real-robot experiments that the identified cost parameters can be used to control sufficiently close dynamics, with performance degrading linearly with the deviations in the dynamics and the identified cost parameters.

Related papers