A Smooth Polynomial Lyapunov Certificate for Convergence of Q-Learning and Its Smooth Variants

arXiv:2404.14442 · cs.LG, cs.AI · Submitted 2024-04-20 · Read on arXiv

cs.LG, cs.AI

Submitted: 2024-04-20

Updated: 2026-09-09

License: http://creativecommons.org/licenses/by/4.0/

The gist: Classical convergence analyses of Q-learning rely on the infinity-norm contraction of Bellman operators, and existing ordinary differential equation (ODE) arguments often use the non-differentiable

Terminology

Abstract

Classical convergence analyses of Q-learning rely on the infinity-norm contraction of Bellman operators, and existing ordinary differential equation (ODE) arguments often use the non-differentiable infinity-norm directly. This paper develops a smooth polynomial Lyapunov-function-based stability certificate for convergence of Q-learning by transferring infinity-norm contraction to a weighted degree- 2p polynomial Lyapunov function induced by a finite 2p-norm. The framework is conceptual and structural: it avoids non-differentiability, handles preconditioned dynamics arising in Q-learning and its variants, and gives a unified stability argument for standard Q-learning and smooth variants based on log-sum-exp (LSE), mellowmax, and Boltzmann softmax operators. For contractive operators, including the max, LSE, and mellowmax cases, the associated ODEs are globally exponentially stable and, under the stated independent and identically distributed (i.i.d.) sampling model, the stochastic approximation iterates converge almost surely. For the Boltzmann operator, which need not be contractive, the same framework yields convergence to an explicit invariant error set around the optimal Q-function. The resulting theory is not intended as a finite-time bound, but as a clean ODE foundation that unifies and simplifies asymptotic analyses of Q-learning and its smooth variants.

Sources

Related papers