Finite-time boundary collision in planar linear quadratic regulator gradient flows

arXiv:2610.00297 · math.OC, cs.SY, eess.SY, math.DS · Submitted 2026-09-26 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: I'm Rosa, and with me are Dev and Taro, guest researcher.

Dev: Today's paper: "Finite-time boundary collision in planar linear quadratic regulator gradient flows".

Rosa: Finite-time boundary collision in planar linear quadratic regulator gradient flows investigates whether an optimization trajectory for LQR can reach the stability boundary in finite time when evaluated from a fixed…

Dev: First, who's behind it and why it matters.

Title and authors: Rosa: The paper is titled "Finite-time boundary collision in planar linear quadratic regulator gradient flows," and it was written by Kang Liu from the School of Future Technology at Xi’an Jiaotong University, China. It immediately tells you the focus is on how LQR optimization trajectories behave when evaluated from a single fixed initial state.

Dev: Kang Liu's work is interesting because it moves beyond just proving convergence in general; it specifically targets that scenario where the cost might diverge near the stability boundary but we are only looking at one specific trajectory.

Taro: I think the authors are setting up a framework to understand when an AI agent, following an LQR policy gradient, will suddenly run into instability rather than smoothly finding its way to a stable operating point.

Rosa: That's right, and the authors claim that for controllable planar systems with one input and positive definite quadratic weights, they can provide an exact representation of the cost which leads to this condition.

Dev: I wonder how their findings translate into practical engineering terms; are we talking about a specific type of system where this finite-time collision is actually a risk in our real-world hardware deployments?

Taro: It’s important because it gives us a precise mathematical boundary for when an autonomous decision-making process becomes fundamentally unstable under the current cost structure.

The paper's summary: Rosa: They summarize by saying that they study whether the Euclidean gradient flow, defined by dK/dτ = −∇KJ(K), can reach the boundary of the stabilizing domain S in finite optimization time Tmax less than infinity.

Dev: That gradient flow description is what I needed; it frames the problem as a dynamic process where we're tracking how the cost changes over time, and they are checking if that process hits a limit too soon.

Taro: The core of the summary is that an exact representation of the accumulated state Gramian and cost function J(K) on a domain where e not equal to zero and a > zero leads to their main result.

Rosa: They establish Proposition three point two which gives an exact formula for the accumulated state Gramian, which is represented as X K = eta / (2Dbay)

z squared + ay kz: / kz k two. That formula seems incredibly specific and technical.

Dev: That specific formula is key because it allows them to represent the Euclidean gradient flow in transformed coordinates as a system of differential equations, k'y' = -M J k/J y. That means they can model the approach direction mathematically.

Taro: And then they introduce a ratio r = k/y and reparametrize time by sigma, leading to vector field equations like r sigma = H0(r) + O(y) and y sigma = G0(r) + O(y). That suggests they are looking for an equilibrium direction.

The paper's improvements: Rosa: One major improvement they highlight is that for the flow to reach the boundary, it has to satisfy beta beta c, where these parameters beta and beta c are derived directly from the specific system data.

Dev: That condition is what gives us a sharp threshold for compact cost sublevels; it means we can determine if we are in a region where convergence is guaranteed, or if we're heading toward that boundary collision based on these calculated parameters.

Taro: The sufficiency proof shows that if beta < beta c, there exists a positively invariant compact rectangle where the extended vector field is smooth, which means trajectories stay confined and converge in finite time T = Z infinityy(sigma)d sigma < infinity.

Rosa: But they also analyzed the critical case where beta = beta c, and they found that even at equality, a specific trapping region is constructed in the curvilinear wedge W, proving that the flow cannot reach y=zero in finite time under those exact conditions.

Dev: That distinction between beta < beta c leading to finite-time convergence and beta = beta c leading to no finite-time collision is a huge piece of information for designing robust control laws.

Conclusion: Rosa: So, summarizing the paper "Finite-time boundary collision in planar linear quadratic regulator gradient flows," the main implication is that they provide a necessary and sufficient condition for predicting whether an LQR optimization trajectory will converge to the optimal Riccati gain or collide with the stability boundary in finite time.

Dev: That prediction capability is what really matters for control engineers because it allows us to pre-emptively design systems that avoid those unstable trajectories entirely, instead of just hoping they stay stable.

Taro: For autonomy, this means we can define explicit parameter regions where we are guaranteed convergence to a stable policy versus basins where finite-time collision is possible when the system misbehaves or parameters shift unexpectedly.

Rosa: I think the sharp threshold for compact cost sublevels is a real asset here, giving us a quantitative certificate that the AI is safely confined to a region where it's converging.

Dev: And tracking that linear vanishing rate of stability margin along colliding trajectories tells us exactly how close we are to that failure point, which lets us implement dynamic safety protocols before any catastrophic failure occurs.

Taro: That precise measurement of instability decay is valuable because it moves our safety analysis from a general statement to a quantitative prediction about the remaining time before an event.

Rosa: It's fascinating how this work connects the abstract mathematics of gradient flows to very concrete, actionable engineering concerns about system stability and failure modes.

Dev: Definitely, so we have this paper, "Finite-time boundary collision in planar linear quadratic regulator gradient flows," which gives us tools to better predict and manage the stability limits of LQR optimization trajectories.

Kang Liu

School of Future Technology, Xi’an Jiaotong University

math.OC, cs.SY, eess.SY, math.DS

Submitted: 2026-09-26

Updated: 2026-09-26

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 92/100

The gist: Finite-time boundary collision in planar linear quadratic regulator gradient flows investigates whether an optimization trajectory for LQR can reach the stability boundary in finite time when

Key concepts

Gradient Flow
This describes how a system moves along the steepest descent path of a cost function. In this context, it models how an LQR control trajectory evolves over time by following the negative gradient of the cost function, trying to minimize that cost.
Stability Boundary
This is the limit where all eigenvalues of the closed-loop system matrix have zero real parts. It represents the edge of a region where a system is guaranteed to be stable; hitting this boundary means reaching a state where stability is lost in finite time.
Riccati Gain (K)
The Riccati gain is the optimal control law derived from solving the LQR problem. The analysis determines if the gradient flow converges toward this specific, optimal gain or if it instead hits the stability boundary before reaching it.

Terminology

Summary

Finite-time boundary collision in planar linear quadratic regulator gradient flows investigates whether an optimization trajectory for LQR can reach the stability boundary in finite time when evaluated from a fixed initial state. This study establishes necessary and sufficient conditions for such collisions, providing a sharp threshold for compact cost sublevels and classifying trajectories as either converging to the optimal Riccati gain or colliding with the stability boundary.

The gist: An exact representation of the cost yields a necessary and sufficient condition for the existence of such a collision, reducing it to a scalar root calculation that determines whether every trajectory either converges to the Riccati gain or collides with the stability boundary in finite optimization time.

Problem Formulation and Collision Criterion

The analysis considers a continuous-time linear system with static state feedback, defining the closed-loop matrix and the stabilizing domain by requiring all eigenvalues of the closed-loop matrix to have strictly negative real parts. The core problem is studying the gradient flow defined by dK/dτ = −∇KJ(K), where J(K) is the cost from a fixed initial state x0, and determining if this flow reaches the boundary of the stabilizing domain S in finite optimization time Tmax < ∞.

The criterion for collision involves several steps:

  1. Identifying the only possible boundary endpoint compatible with bounded cost.

  2. Determining whether the gradient flow can approach that endpoint from within S.

This is achieved by transforming system data using an orthogonal matrix U to simplify the dynamics, leading to a closed-loop matrix structure where a finite-time collision must have a specific form when e ≠ 0 and a > 0. The criterion is summarized by Theorem 2.2, which states that for the flow to reach the boundary, it must satisfy β ≤ βc, where β and βc are derived from system data.

Exact Cost and Induced Metric

The paper establishes an exact rational representation of the accumulated state Gramian (XK) and the cost function J(K) on a domain where e ≠ 0 and a > 0. The key result is Proposition 3.2, which provides an exact formula for XK in terms of gain coordinates (k, y):

XK = η / (2Dbay) [z2 + ay kz] / [kz k2].

The Euclidean gradient flow is then represented by a system of differential equations in the transformed coordinates: k'y' = −M Jk/Jy. The metric M is derived from the Jacobian of the gain transformation, and its inverse appears in this representation. This formulation allows for a smooth extension of the dynamics to y = 0 for finite r, which is crucial for analyzing approach directions.

Direction of Approach and Radial Motion

To analyze trajectories approaching the boundary (k, y) = (0, 0), the analysis introduces a ratio r = k/y and reparametrizes time by σ. This leads to a vector field equation: rσ = H0(r) + O(y), yσ = G0(r) + O(y).

The existence of an equilibrium direction requires solving the equation (17): (1 + γr)G0(r)e2C − r2 = γP(r). The analysis shows that for a root r∗ of P, the radial coefficient is determined by G0(r∗) = e2Cr2 / (1 + γr∗), and the condition for a negative radial speed is θ∗ = −γr > 1.

Sufficiency, Including the Critical Equality

The sufficiency proof demonstrates that if β < βc, there exists a positively invariant compact rectangle in the (u, y) coordinates where the extended vector field is smooth. This invariance guarantees that every interior trajectory satisfies 0 < y(σ) ≤ y(0)e−gσ, leading to a finite time T = Z∞0y(σ)dσ < ∞.

Crucially, at the equality case (β = βc), an explicit trapping region is constructed in the curvilinear wedge W. This construction proves that even at equality, the flow cannot reach y=0 in finite time because the differential inequalities imply that u and y do not reach zero for any finite σ.

Global Dichotomy and Parameter Classification

Specializing to a specific family of systems (25), Theorem 4.1 provides a global dichotomy: every stabilizing initial gain has exactly one outcome: either it converges to the unique stabilizing Riccati gain K⋆, or it reaches the stability boundary in finite optimization time T, with J(K(τ)) ↓ 1/2 as τ ↑ T.

The collision region is defined by the parameter inequality E(h, v) ≥ 0. The analysis shows that for this specific family, a non-empty open collision basin exists when v ≥ 25/2 and h−v ≤ h ≤ h+(v).

Improvements for AI systems

As a fastidious researcher, I have analyzed this paper, Finite-time boundary collision in planar linear quadratic regulator gradient flows, and identified several high-impact applications for improving AI systems.

The core insight of the paper is that policy gradient methods for optimal control (like those used to train reinforcement learning agents) can reach stability boundaries in finite time, which has significant implications for training stability and convergence guarantees.

Here are the specific improvements and capabilities derived from this research:


) Specific Improvements & Enhanced AI Capabilities:


  1. Enhanced Stability Guarantees for Policy Gradient Methods (PPO/TRPO/SAC):

  2. The paper provides a necessary and sufficient condition for when a policy gradient trajectory converges to the stability boundary in finite time. This allows researchers to move beyond the assumption that cost descent prevents reaching instability.

  3. Specifically, by calculating the critical threshold parameter region (Theorem 4.1), AI developers can precisely map out which initial policies (gains) are guaranteed to converge to a stable, optimal policy versus those that are destined for finite-time boundary collision (instability).


  4. Finite-Time Convergence Certificates:

  5. The existence of a sharp threshold for compact cost sublevels provides a quantitative certificate of convergence. An AI system can monitor its current cost and Gramian properties in real-time to determine if it is confined to a compact, stable region (where convergence to the optimal policy is guaranteed) or if it is on a trajectory leading toward instability.


  6. Robustness Analysis Against Instability:

  7. The analysis shows that along every colliding trajectory, the stability margin and the smallest eigenvalue of the accumulated state Gramian vanish linearly in time, even though the Gramian remains positive definite before collision. This allows for a precise understanding of how near-miss scenarios (where performance degrades slowly) behave dynamically.

  8. AI agents can be trained with explicit safeguards that monitor this linear vanishing rate; if the rate exceeds a certain bound, it signals an impending finite-time failure, allowing for preemptive intervention or trajectory modification before catastrophic failure occurs.


  9. Parameter-Specific Training Regimes:

  10. The paper provides explicit algebraic parameter regions (Section 4) for specific system configurations (e.g., the triangular family). This allows AI engineers to design training regimes tailored to the system's dynamics, knowing exactly which coupling strengths or weights will result in convergence versus collision, rather than relying on broad empirical testing.

  11. This enables the creation of collision basins (regions where finite-time collision is possible) and convergence regions (where convergence to the Riccati gain is guaranteed), allowing for targeted exploration of the policy space during training.

13.---

) What the Improved AI System Can Do:

The improved AI system—a robust, formally verified Policy Gradient Optimizer—can perform:


  1. Determine Convergence vs. Collision Status: The system can definitively classify its current state as either converging to the optimal stabilizing policy or heading toward a finite-time stability boundary collision based on its real-time cost and Gramian measurements against the derived necessary and sufficient conditions (Theorem 2.2).


  2. Implement Dynamic Safety Protocols: If the system detects that its trajectory is entering a region where the convergence guarantee is lost (i.e., approaching the critical parameter threshold), it can trigger dynamic safety protocols, such as increasing exploration noise or modifying control inputs, to steer the policy away from a collision basin and toward a guaranteed compact sublevel.


  3. Quantify Stability Margin Decay: The system can precisely track the linear rate at which its stability margin vanishes near the boundary (as derived in Proposition 3.3). This provides a quantitative measure of how close it is to failure, allowing for much finer-grained control than simple binary stable/unstable flags.

8.---

  1. System Design and Training Optimization: AI researchers can use the explicit parameter regions (e.g., Theorem 4.1) to design training environments that either force convergence or deliberately create a collision basin for study, optimizing the exploration strategy of the agent based on theoretical guarantees rather than just empirical observation.

Sources

Related papers