Algorithmic Optimality Guarantees for Nonsmooth H infinity Output-Feedback Policy Search
math.OC, cs.LG, cs.SY, eess.SY
Submitted: 2026-09-05
Updated: 2026-09-05
Comments: Appeared at The 65th IEEE Conference on Decision and Control (CDC), 2026
License: http://creativecommons.org/licenses/by/4.0/
The gist: We study continuous-time full-order dynamic output-feedback H infinity policy search, a nonconvex and nonsmooth problem.
Terminology
Abstract
We study continuous-time full-order dynamic output-feedback H infinity policy search, a nonconvex and nonsmooth problem. Direct policy search is a central paradigm in reinforcement learning and continuous control, but rigorous guarantees remain scarce in robust output-feedback settings. The H infinity problem is a canonical benchmark because it captures disturbance attenuation and robustness while exposing the hard nonsmooth geometry of policy-space optimization. We prove that on the exact identity-gauge slice of the extended convex lift, epsilon-stationarity yields O(epsilon) -suboptimality on compact exact slices, which in turn yields convergence-rate guarantees for nonsmooth policy-search methods. This result addresses the finite-time optimality-gap question raised by Guo and Hu [2022] in the more general dynamic output-feedback H infinity policy-search setting. We further use the established value equivalence supplied by extended convex lifting to formulate a nonstrict-feasibility bisection method with one final strict-feasibility recovery step, yielding an explicit epsilon-optimal stabilizing controller. These results provide a quantitative and algorithmic strengthening of prior qualitative optimality theory for nonsmooth H infinity policy search.
Related papers
- Lions and Muons: Optimization via Stochastic Frank-Wolfe under Heavy-Tailed Noise
- Adam-HNAG: A Convergent Reformulation of Adam with Accelerated Rate
- Incremental Learning in Mirror Flows
- Online Control via Counterfactual Tracking
- Asynchronous Replanning in Two Population Linear Quadratic Mean Field Games: Information Requirements and Stability
- Petrov-Galerkin operator inference with application to stability-encouraging identification