Predefined-Time Integral Reinforcement Learning for Unknown Nonlinear Systems via Inverse-Optimal Design

arXiv:2610.12103 · eess.SY, cs.SY · Submitted 2026-10-08 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: Today's paper: "Predefined-Time Integral Reinforcement Learning for Unknown Nonlinear Systems via Inverse-Optimal Design".

Dev: The gist This paper develops a new predefined-time integral reinforcement learning framework for optimal control of unknown nonlinear systems via inverse-optimal design.

Rosa: First, who's behind it and why it matters.

Paper summary: Rosa: So, to wrap up what we just discussed about the "Predefined-Time Integral Reinforcement Learning for Unknown Nonlinear Systems via Inverse-Optimal Design," the paper is proposing a new framework for controlling nonlinear systems where you don't know the exact dynamics.

Dev: The central idea is to use a predefined-time integral reinforcement learning method, which means you set a target convergence time upfront, and then design the control objective around that specific time constraint.

Taro: It claims they can do this by first approximating the unknown drift using an RBF neural network and an online identification law to get (x), and then incorporating that into their integral reinforcement learning problem.

Rosa: They introduce a specific objective function J f for the identifier that balances fitting the current data with some stored integral predictions, and then they use this reconstructed drift (x) to build an inverse-optimal control policy.

Dev: The main claim is that this approach allows for the construction of a feedback law whose optimal policy inherits a predefined-time stabilization property, which is achieved by selecting a desired value function and prescribing its Lyapunov decay behavior in advance.

Taro: They are essentially showing how to decouple the problem so that even though f(x) is unknown during the critic learning, the controller can still be designed optimally based on what it has learned from integral data.

Rosa: And they prove this works by establishing a certified deadline T p, which is derived from flushing information from a moving window and bounding the remaining composite convergence time using T E and T TX.

Dev: So, the overall message is that you can achieve optimal control for unknown nonlinear systems with a guaranteed convergence time bound if you use this specific combination of drift identification, integral reinforcement learning, and inverse-optimal design.

Taro: What this means for autonomy research is that we can build systems that are not only autonomous but also predictable in terms of their stabilization speed when faced with uncertain physical dynamics.

Conclusion: Rosa: Looking at the title, "Predefined-Time Integral Reinforcement Learning for Unknown Nonlinear Systems via Inverse-Optimal Design," it sounds like this paper is tackling a really tough problem in control where the dynamics are nonlinear and completely unknown.

Dev: And what it does is move beyond just finding *a* working controller to designing one that works within a specific time budget you set beforehand, which is what the predefined-time aspect delivers.

Taro: The authors, Tien Dat Vua and their team, are showing how to combine neural network identification with reinforcement learning techniques in a way that respects these hard real-time constraints.

Rosa: The implication for us on the ground is that for things like autonomous vehicles or complex robots operating outside of perfectly controlled labs, we can finally have a mechanism where we know exactly when the system will settle down to its stable state.

Dev: It’s about moving from reactive control, where you hope it converges eventually, to proactive control, where you engineer the convergence deadline into the learning process itself.

Taro: So, in simple terms for someone just listening here, this paper shows a way to build a learning system that doesn't just learn how to behave, but learns how fast it can reliably get there.

Rosa: That's right. It’s about making the learning process time-aware so we have more confidence in deploying these types of systems in the real world where timing matters a lot.

Tien Dat Vua

Ho Chi Minh City University of Technology (HCMUT) · Vietnam National University Ho Chi Minh City (VNU-HCM)

eess.SY, cs.SY

Submitted: 2026-10-08

Updated: 2026-10-08

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

The gist: The gist This paper develops a new predefined-time integral reinforcement learning framework for optimal control of unknown nonlinear systems via inverse-optimal design.

Key concepts

System Identification
This step reconstructs the unknown nonlinear drift function f(x) of the system using a radial basis function neural network and an online learning law. This allows the controller to estimate how the system moves based on observed data.
Predefined-Time Stability
This concept ensures that for any starting state, the system is guaranteed to reach a desired state (like zero) within a specific, predetermined time bound (Ts). This is achieved by designing a Lyapunov function whose derivative guarantees this time constraint.
Inverse-Optimal Design
Instead of learning the control policy directly from scratch, this approach uses predefined stability results to select the optimal value function and Lyapunov decay rate. This guides the learning process toward an optimal control law that respects the desired convergence time.
Critic-Only Integral Reinforcement Learning
This method learns both the value function and control policy without knowing the exact system dynamics beforehand. It updates its estimates using current integral data and a stored replay stack, ensuring practical stability even without perfect information.

Terminology

Summary

The gist This paper develops a new predefined-time integral reinforcement learning framework for optimal control of unknown nonlinear systems via inverse-optimal design.

Problem Formulation and System Identification

Consider the nonlinear control-affine system x˙ = f(x) + g(x)u, (1) where x ∈ R n is the system state, u ∈ R m is the control input,

The unknown drift f(x) is first reconstructed from measured data using a radial basis function (RBF) neural network and an online identification law for updating its weights The identified drift is represented by ˆf(x) = Wˆ ⊤ f σf(x), where Wˆ f denotes the estimated weight matrix associated with the prescribed basis vector σf(x)

The integral composite concurrent-learning mechanism combines current moving-window information with informative stored integral data to update the drift weights The objective function for the identifier is defined as Jf:= 1/2∥ef(t)∥2 + kf/2 PMf −1l=0∥ef,l(t)∥2, where ef(t) and ef,l(t) are the current and stored integral prediction errors

Predefined-Time Reinforcement Learning Optimal Control

The normalized predefined-time comparison result in Theorem 1 provides a direct mechanism for assigning a desired convergence-time bound in advance,

The designer selects a continuously differentiable positive-definite function Vd: omega → R≥0 with Vd(0) = 0, and prescribes its desired closedloop evolution as V˙d(x) = −γp,q,r Ts h αVp d(x) + βVq d(x) ir According to Theorem 1, any feedback law satisfying (22) renders the origin predefined-time stable and guarantees T(x0) ≤ Ts, ∀ x0 ∈ R n

The optimal control policy is constructed using inverse-optimal design by selecting a desired value function and a prescribed Lyapunov decay, and then constructing a running cost for which the resulting feedback law is optimal The resulting optimal feedback satisfies V˙d(x) ≤ −γp,q,r Ts h αVp d(x) + βVq d(x) ir

Critic-Only Integral Reinforcement Learning

In the present IRL formulation, f(x) is still regarded as unknown, and the optimal value function and control policy are learned directly from integral data without requiring the exact drift dynamics,

The critic estimate is Vˆ (x) = Wˆ ⊤ϕ(x), where ϕ(x) ∈ R N c, Wˆ ∈ R N c, and ε(x) denotes the approximation error The critic-generated policy is uˆ(x,Wˆ) = − 1/2 R−1 h L⊤ 1(x) + g⊤ (x)(∇ϕ(x))⊤Wˆ i

The critic update law is formulated using current integral Bellman data together with a finite informative replay stack, thereby guaranteeing practical predefined-time critic learning without requiring persistent excitation throughout the closed-loop operation The scaled gradient flow is defined as ˙Wˆ = −αW Γ∇Wˆ JW, where αW > 0 and Γ = Γ⊤≻ 0

Predefined-Time Closed-Loop Stability Analysis

The certified deadline is Tp = TE+TW +T+TX: T flushes prelearning information from the moving window, and TX bounds the remaining composite convergence time,

Theorem 4 establishes that J(t) enters no later than Tp the forward-invariant set BJ:= n J ≥ 0: J ≤ J¯ o, C1J¯ µ + C2J¯ν = ∆J/θJ Moreover,∥x(t)∥ ≤ rx:= s 1/v max (2LV˙ T, 2J¯ T), t ≥ Tp

The final result shows that col of x and W˜ is practically predefined-time stable If ε¯B(t) ≡ 0, ε¯B, j = 0 for j = 1,..., M, and ε¯g = 0, then x(t) = 0 and W˜ (t) = 0 for all t ≥ Tp

Design Rule

**Corollary 2 (Overall-deadline-based selection of Ts). Let Tp > TE + T be a designer-assigned overall convergence deadline, and choose χ ∈ (0, 1).

Improvements for AI systems

  1. A unified predefined-time integral reinforcement learning framework is developed for unknown nonlinear systems by integrating online drift identification, integral concurrent learning, inverse-optimal design, and critic-based reinforcement learning within a single architecture. This enables the prescribed-time optimal-control construction to be carried out without prior knowledge of the nonlinear drift.

  2. A new critic update law is developed using current integral Bellman data together with a finite informative replay stack, thereby guaranteeing practical predefined-time critic learning without requiring persistent excitation throughout the closed-loop operation. This addresses the limitation that the corresponding Bellman residual still contains explicit system-dynamics terms in prior work.

  3. The framework allows for the construction of an optimal feedback law by defining a cost function where the desired convergence time is introduced directly into the control objective, which ensures the optimal policy inherits the predefined-time stabilization property.

  4. The system can achieve practical predefined-time closed-loop stability, as shown by Theorem 4, which guarantees that the state trajectory enters a set whose size is bounded by a function of the designer's parameters and constraints: ∥x(t)∥ ≤ rx:= s/v max (2LV˙ T, 2J¯/T), t ≥ Tp.

  5. The system can operate under aggressive timing requirements, as illustrated by the simulation results where the worstcase state norm is maxi ∥xi(Tp)∥ = 2.7008×10−3 for a prescribed deadline of one second, demonstrating that the bound is achievable but requires a larger transient control action.

  6. The system can operate under relaxed timing specifications, as shown by the results for Tp = 100 s where the state trajectories converge long before the prescribed deadline, indicating that the current convergence-time bound remains conservative.

Related papers