Forward Trajectory Steering for Hamilton-Jacobi Reachability Analysis

arXiv:2608.11480 · eess.SY, cs.LG, cs.SY · Submitted 2026-08-11 · Read on arXiv

Sungje Park, Stephen Tu

University of Southern California

eess.SY, cs.LG, cs.SY

Submitted: 2026-08-11

Updated: 2026-08-13

Comments: IEEE CDC 2026

Code: https://github.com/sungje-park/steer2reach

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 75/100

The gist: This paper introduces STEER2REACH (S2R), a physics-informed neural network (PINNs)-based solver for Hamilton-Jacobi (HJ) reachability analysis.

Terminology

Summary

This paper introduces STEER2REACH (S2R), a physics-informed neural network (PINNs)-based solver for Hamilton-Jacobi (HJ) reachability analysis. The work addresses the computational bottleneck of solving Hamilton-Jacobi-Isaacs variational inequality (HJI-VI) PDEs in high dimensions, which traditionally suffer from the curse of dimensionality. The authors propose a lightweight adaptive collocation sampling strategy that requires minimal modification on top of standard PINNs training.

HJ reachability analysis provides a mathematically rigorous framework for safe control of dynamical systems by computing a safety value function that identifies safe operation regimes. The value function V(x,t) is computed as the viscosity solution to the HJI-VI:

RV(x,t) = min ∂tV(x,t) + H(x,t), l(x) − V(x,t)

where H(x,t) = max min ⟨∇xV(x,t), f(x,u,d)⟩, with u representing control and d representing disturbance.

The Backward Reachable Tube (BRT) is defined as B = x V(x,0) ≤ 0, containing initial states from which the system will inevitably enter a failure set L under worst-case disturbance.

The authors note that while PINNs have emerged as alternatives to classical mesh-based solvers, direct application of PINNs to HJI-VI PDEs does not necessarily yield high-fidelity value functions. Existing state-of-the-art (SoTA) PINNs-based reachability solvers incorporate complex techniques including:

  • Periodic activation functions

  • Exact boundary constraints

  • Time-based curriculum learning

  • Linear semi-supervision

  • MPC-guided sampling with multi-stage training pipelines

The central research question is: Can one design a simple adaptive sampling strategy for PINNs-based HJ reachability analysis that achieves near SoTA performance on high-dimensional problems?

S2R draws inspiration from backward stochastic differential equation (BSDE) methods for solving PDEs. The key innovation is constructing an adaptive sampling distribution by:

  1. Using the current value function to compute optimal control and disturbance signals that guide trajectory rollouts

  2. Injecting stochasticity around these trajectories by modeling the system as an SDE with a tunable noise parameter

The method extends deterministic dynamics to an SDE: dX = f(X,u,d)dτ + σ(X,u,d)dW, where W denotes standard Brownian motion.

The optimal control and disturbance are computed from the current value function Vθ:

  • uθ(x,t) = argmax min ⟨∇Vθ(x,t), f(x,u,d)⟩

  • dθ(x,t) = argmin ⟨∇Vθ(x,t), f(x,uθ(x,t),d)⟩

The S2R loss is defined as: LS2R(θ; θsamp, λ) = LVI(θ; µθsamp) + λLbnd(θ; µ′θsamp), where µθ and µ′θ are measures induced by the forward SDE trajectories.

The optimization uses repeated gradient descent (RGD), where the sampling distribution changes as a function of the current parameter value, a framework known as performative prediction.

The authors propose a parameterization that enforces two constraints simultaneously:

Vθ(x,t) = l(x) − (T − t)ρ(ϕθ(x,t))

where ρ is a fixed non-learnable function with non-negative outputs (e.g., quadratic ρ(x) = x2). This enforces Vθ(x,t) ≤ l(x) for all x and removes the need for a boundary loss term.

The authors evaluate S2R against standard PINNs, RAD-PINNs (residual-based adaptive distribution), and MPC-DeepReach across five benchmarks: 2D Vertical Drone, 3D Pursuit-Evade, 7D F1Tenth, 13D Quadrotor, and 40D Publisher-Subscriber.

Main findings from Table I:

  1. 2D Vertical Drone: S2R achieves the best RL2 error (0.0497 vs. 0.1020 for MPC-DeepReach) with comparable safety metrics (IOU: 0.9745 vs. 0.9688)

  2. 3D Pursuit-Evade: S2R achieves the lowest RL2 (0.0271 vs. 0.0360 for MPC-DeepReach) with competitive IOU (0.9854 vs. 0.9910)

  3. 7D F1Tenth: MPC-DeepReach outperforms S2R (IOU: 0.9603 vs. 0.7364), attributed to hybrid dynamics sensitivity

  4. 13D Quadrotor: MPC-DeepReach has slight advantage on safety metrics, but S2R is competitive

  5. 40D Publisher-Subscriber: S2R achieves the best performance across all metrics, with precision 0.9979, IOU 0.9878, and RL2 0.0983 (vs. 0.5676 for MPC-DeepReach)

The authors note: "S2R is able to be competitive with and can match the performance of MPC-DeepReach on most problems... while achieving a lower RL2 (ground truth residual error) by a factor of 1.3× for the pursuit-evade, 2× for the vertical drone, and 5.7× for the publisher-subscriber problem."

The performance gap on F1Tenth is attributed to hybrid dynamics which are piecewise nonlinear and numerically sensitive with a narrow operating region. The high-speed dynamics introduce terms inversely proportional to v and v2 making the yaw and slip angle equations highly sensitive to discretization error. The authors tried clipping, bounding, and filtering strategies but found these introduced a critical sampling bias.

In experiments varying the publisher-subscriber problem from n=2 to n=500 dimensions, the performance of S2R in RL2 scales significantly better compared to the PINNs baseline as dimensions increases, demonstrating effective sampling of relevant state space regions even in high dimensions.

  1. Constraint functions: Testing quadratic, softplus, swish, and elu constraints showed hard enforcement of the inequality constraint can lead to improved performance... for both PINNs and S2R, though there is not a clear trend as to which choice of ρ is the best.

  2. Noise parameter σ: RL2 error can be improved by tuning the noise, illustrating that the exploration induced by the forward trajectory SDE can be helpful for learning, though the best choice is heavily problem dependent.

  3. Discretization time ∆t: Smaller ∆t does not necessarily lead to lower error despite more accurate SDE integration.

The authors demonstrate that S2R achieves competitive safety performance and lower relative L2 error compared to SoTA solvers, while being much simpler to implement and not requiring multi-stage training pipelines or auxiliary semi-supervision. The method introduces only two new hyperparameters: noise injection σ and integration timestep ∆t.

Future directions include designing reflected BSDE solvers that yield even higher quality safety value functions than S2R and applying S2R to visuomotor control setting for enabling scalable and accurate latent reachability analysis.

Improvements for AI systems

Based on the paper, here are the specific improvements I can make to AI systems and what the improved systems can do:

Improvement: Replace static or uniform collocation point sampling with a trajectory-guided adaptive sampling mechanism that uses the current model's gradient information to generate optimal control/disturbance signals, then injects stochastic noise around those trajectories.

Capability: The improved system can solve high-dimensional PDEs (up to 500 dimensions) with significantly lower residual error—5.7× better than existing methods on 40D problems—while requiring only two additional hyperparameters (noise σ and timestep Δt) instead of complex multi-stage training pipelines.

Improvement: Enforce inequality constraints (V(x,t) ≤ l(x)) directly through the network architecture using a non-learnable positive function (e.g., quadratic), rather than adding soft penalty terms to the loss.

Improvement: Treat the sampling distribution as a function of the current model parameters and use repeated gradient descent to handle the distribution shift that occurs as the model improves.

Improvement: Model deterministic dynamics as stochastic differential equations with tunable noise to inject controlled exploration around optimal trajectories during training.

Improvement: Use the S2R framework to compute backward reachable tubes and safety value functions in high-dimensional state spaces without mesh-based discretization.

Improvement: Eliminate the need for time-based curriculum learning or multi-stage training by using trajectory-guided sampling that naturally focuses on relevant regions of the state-time space.

  • Solve high-dimensional safety-critical control problems (up to 500D) that were previously intractable, enabling safe operation of complex autonomous systems in real-world scenarios.

  • Provide certified safety guarantees through accurate value function approximation, with lower residual error than existing neural solvers, making it suitable for safety-critical applications like drone navigation, autonomous racing, and multi-agent coordination.

  • Adapt to new problem domains quickly with minimal hyperparameter tuning (only σ and Δt), reducing the expertise required to deploy neural PDE solvers in new applications.

  • Handle hybrid dynamics (piecewise nonlinear systems) with appropriate sensitivity analysis, though the system can identify when such dynamics require specialized treatment.

  • Generate high-quality training data on-the-fly by simulating optimal trajectories, eliminating the need for pre-computed datasets or expensive mesh generation in high dimensions.

  • Enable latent reachability analysis for visuomotor control, potentially extending safety verification to systems with image-based observations, as suggested by the paper's future directions.

Sources

Related papers