Forward Trajectory Steering for Hamilton-Jacobi Reachability Analysis
Sungje Park, Stephen Tu
University of Southern California
eess.SY, cs.LG, cs.SY
Submitted: 2026-08-11
Updated: 2026-08-13
Comments: IEEE CDC 2026
Code: https://github.com/sungje-park/steer2reach
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 75/100
The gist: This paper introduces STEER2REACH (S2R), a physics-informed neural network (PINNs)-based solver for Hamilton-Jacobi (HJ) reachability analysis.
Terminology
Summary
This paper introduces STEER2REACH (S2R), a physics-informed neural network (PINNs)-based solver for Hamilton-Jacobi (HJ) reachability analysis. The work addresses the computational bottleneck of solving Hamilton-Jacobi-Isaacs variational inequality (HJI-VI) PDEs in high dimensions, which traditionally suffer from the curse of dimensionality. The authors propose a lightweight adaptive collocation sampling strategy that requires minimal modification on top of standard PINNs training.
HJ reachability analysis provides a mathematically rigorous framework for safe control of dynamical systems
by computing a safety value function that identifies safe operation regimes. The value function V(x,t) is computed as the viscosity solution to the HJI-VI:
RV(x,t) = min ∂tV(x,t) + H(x,t), l(x) − V(x,t)
where H(x,t) = max min ⟨∇xV(x,t), f(x,u,d)⟩, with u representing control and d representing disturbance.
The Backward Reachable Tube (BRT) is defined as B = x V(x,0) ≤ 0, containing initial states from which the system will inevitably enter a failure set L under worst-case disturbance.
The authors note that while PINNs have emerged as alternatives to classical mesh-based solvers, direct application of PINNs to HJI-VI PDEs does not necessarily yield high-fidelity value functions.
Existing state-of-the-art (SoTA) PINNs-based reachability solvers incorporate complex techniques including:
-
Periodic activation functions
-
Exact boundary constraints
-
Time-based curriculum learning
-
Linear semi-supervision
-
MPC-guided sampling with multi-stage training pipelines
The central research question is: Can one design a simple adaptive sampling strategy for PINNs-based HJ reachability analysis that achieves near SoTA performance on high-dimensional problems?
S2R draws inspiration from backward stochastic differential equation (BSDE) methods for solving PDEs. The key innovation is constructing an adaptive sampling distribution by:
-
Using the current value function to compute optimal control and disturbance signals that guide trajectory rollouts
-
Injecting stochasticity around these trajectories by modeling the system as an SDE with a tunable noise parameter
The method extends deterministic dynamics to an SDE: dX = f(X,u,d)dτ + σ(X,u,d)dW, where W denotes standard Brownian motion.
The optimal control and disturbance are computed from the current value function Vθ:
-
uθ(x,t) = argmax min ⟨∇Vθ(x,t), f(x,u,d)⟩
-
dθ(x,t) = argmin ⟨∇Vθ(x,t), f(x,uθ(x,t),d)⟩
The S2R loss is defined as: LS2R(θ; θsamp, λ) = LVI(θ; µθsamp) + λLbnd(θ; µ′θsamp), where µθ and µ′θ are measures induced by the forward SDE trajectories.
The optimization uses repeated gradient descent (RGD), where the sampling distribution changes as a function of the current parameter value, a framework known as performative prediction.
The authors propose a parameterization that enforces two constraints simultaneously:
Vθ(x,t) = l(x) − (T − t)ρ(ϕθ(x,t))
where ρ is a fixed non-learnable function with non-negative outputs (e.g., quadratic ρ(x) = x2). This enforces Vθ(x,t) ≤ l(x) for all x and removes the need for a boundary loss term.
The authors evaluate S2R against standard PINNs, RAD-PINNs (residual-based adaptive distribution), and MPC-DeepReach across five benchmarks: 2D Vertical Drone, 3D Pursuit-Evade, 7D F1Tenth, 13D Quadrotor, and 40D Publisher-Subscriber.
Main findings from Table I:
-
2D Vertical Drone: S2R achieves the best RL2 error (0.0497 vs. 0.1020 for MPC-DeepReach) with comparable safety metrics (IOU: 0.9745 vs. 0.9688)
-
3D Pursuit-Evade: S2R achieves the lowest RL2 (0.0271 vs. 0.0360 for MPC-DeepReach) with competitive IOU (0.9854 vs. 0.9910)
-
7D F1Tenth: MPC-DeepReach outperforms S2R (IOU: 0.9603 vs. 0.7364), attributed to hybrid dynamics sensitivity
-
13D Quadrotor: MPC-DeepReach has slight advantage on safety metrics, but S2R is competitive
-
40D Publisher-Subscriber: S2R achieves the best performance across all metrics, with precision 0.9979, IOU 0.9878, and RL2 0.0983 (vs. 0.5676 for MPC-DeepReach)
The authors note: "S2R is able to be competitive with and can match the performance of MPC-DeepReach on most problems... while achieving a lower RL2 (ground truth residual error) by a factor of 1.3× for the pursuit-evade, 2× for the vertical drone, and 5.7× for the publisher-subscriber problem."
The performance gap on F1Tenth is attributed to hybrid dynamics which are piecewise nonlinear and numerically sensitive with a narrow operating region.
The high-speed dynamics introduce terms inversely proportional to v and v2 making the yaw and slip angle equations highly sensitive to discretization error.
The authors tried clipping, bounding, and filtering strategies but found these introduced a critical sampling bias.
In experiments varying the publisher-subscriber problem from n=2 to n=500 dimensions, the performance of S2R in RL2 scales significantly better compared to the PINNs baseline as dimensions increases,
demonstrating effective sampling of relevant state space regions even in high dimensions.
-
Constraint functions: Testing quadratic, softplus, swish, and elu constraints showed
hard enforcement of the inequality constraint can lead to improved performance... for both PINNs and S2R,
thoughthere is not a clear trend as to which choice of ρ is the best.
-
Noise parameter σ:
RL2 error can be improved by tuning the noise, illustrating that the exploration induced by the forward trajectory SDE can be helpful for learning,
though the best choice isheavily problem dependent.
-
Discretization time ∆t:
Smaller ∆t does not necessarily lead to lower error despite more accurate SDE integration.
The authors demonstrate that S2R achieves competitive safety performance and lower relative L2 error compared to SoTA solvers, while being much simpler to implement and not requiring multi-stage training pipelines or auxiliary semi-supervision.
The method introduces only two new hyperparameters: noise injection σ and integration timestep ∆t.
Future directions include designing reflected BSDE solvers that yield even higher quality safety value functions than S2R
and applying S2R to visuomotor control setting for enabling scalable and accurate latent reachability analysis.
Improvements for AI systems
Based on the paper, here are the specific improvements I can make to AI systems and what the improved systems can do:
Improvement: Replace static or uniform collocation point sampling with a trajectory-guided adaptive sampling mechanism that uses the current model's gradient information to generate optimal control/disturbance signals, then injects stochastic noise around those trajectories.
Capability: The improved system can solve high-dimensional PDEs (up to 500 dimensions) with significantly lower residual error—5.7× better than existing methods on 40D problems—while requiring only two additional hyperparameters (noise σ and timestep Δt) instead of complex multi-stage training pipelines.
Improvement: Enforce inequality constraints (V(x,t) ≤ l(x)) directly through the network architecture using a non-learnable positive function (e.g., quadratic), rather than adding soft penalty terms to the loss.
Improvement: Treat the sampling distribution as a function of the current model parameters and use repeated gradient descent to handle the distribution shift that occurs as the model improves.
Improvement: Model deterministic dynamics as stochastic differential equations with tunable noise to inject controlled exploration around optimal trajectories during training.
Improvement: Use the S2R framework to compute backward reachable tubes and safety value functions in high-dimensional state spaces without mesh-based discretization.
Improvement: Eliminate the need for time-based curriculum learning or multi-stage training by using trajectory-guided sampling that naturally focuses on relevant regions of the state-time space.
-
Solve high-dimensional safety-critical control problems (up to 500D) that were previously intractable, enabling safe operation of complex autonomous systems in real-world scenarios.
-
Provide certified safety guarantees through accurate value function approximation, with lower residual error than existing neural solvers, making it suitable for safety-critical applications like drone navigation, autonomous racing, and multi-agent coordination.
-
Adapt to new problem domains quickly with minimal hyperparameter tuning (only σ and Δt), reducing the expertise required to deploy neural PDE solvers in new applications.
-
Handle hybrid dynamics (piecewise nonlinear systems) with appropriate sensitivity analysis, though the system can identify when such dynamics require specialized treatment.
-
Generate high-quality training data on-the-fly by simulating optimal trajectories, eliminating the need for pre-computed datasets or expensive mesh generation in high dimensions.
-
Enable latent reachability analysis for visuomotor control, potentially extending safety verification to systems with image-based observations, as suggested by the paper's future directions.
Sources
- MADR: MPC-guided Adversarial DeepReach
- Bridging Model Predictive Control and Deep Learning for Scalable Reachability Analysis
- Predictive Limitations of Physics-Informed Neural Networks in Vortex Shedding
- Numerical simulation of BSDEs using empirical regression methods: theory and practice
- A Forward Reachability Perspective on Control Barrier Functions and Discount Factors in Reachability Analysis
- Generalizing Safety Beyond Collision-Avoidance via Latent-Space Reachability Analysis
Related papers
- One Request, Multiple Experts: LLM Orchestrates Domain Specific Models via Adaptive Task Routing
- A Geometric Decision Procedure for STL Feasibility and Repair
- Submodular Multi-Agent Policy Learning for Online Distributed Task Allocation in Open Multi-Agent Systems
- Policy-Level Recursive Self-Improvement for Embodied AI with a Criticality World Model
- Minimal Experiments for Robust Stabilization: Information, Spectral Geometry, and Duration
- Decentralized Power-Optimal Coordination for Spacecraft Swarms Using Time-Varying Magnetorquer Actuation