Computationally efficient safe exploration in reinforcement learning

arXiv:2609.22919 · cs.LG, cs.RO, cs.SY, eess.SY · Submitted 2026-09-19 · Read on arXiv

cs.LG, cs.RO, cs.SY, eess.SY

Submitted: 2026-09-19

Updated: 2026-09-19

Comments: 8 pages

License: http://creativecommons.org/licenses/by-sa/4.0/

The gist: Reinforcement learning in real-life applications requires safety guarantees during exploration.

Terminology

Abstract

Reinforcement learning in real-life applications requires safety guarantees during exploration. Typical reinforcement learning algorithms do not provide such guarantees, and many modifications that do rely on Gaussian processes (GPs), which have a large computational cost. We propose a computationally lightweight algorithm based on the Nadaraya-Watson estimator that safely explores and optimizes constrained Markov decision processes (MDPs). Our algorithm, CoLSafe-MDP, uses an estimator that scales in constant-time with bounds on the estimates, a significant improvement from its GP-based counterparts that scale cubically with the number of data points. We then evaluate its performance in a grid-based environment and on observational Martian terrain data.

Related papers