The Bias of Nonlinear Two-Time-scale Stochastic Approximation under Constant Step-Sizes

arXiv:2609.20409 · cs.LG, math.OC, stat.ML · Submitted 2026-09-17 · Read on arXiv

cs.LG, math.OC, stat.ML

Submitted: 2026-09-17

Updated: 2026-09-17

License: http://creativecommons.org/licenses/by/4.0/

The gist: Two-timescale stochastic approximation (TTSA) is a fundamental tool for analyzing coupled iterative algorithms in reinforcement learning, optimization, and stochastic control.

Terminology

Abstract

Two-timescale stochastic approximation (TTSA) is a fundamental tool for analyzing coupled iterative algorithms in reinforcement learning, optimization, and stochastic control. However, finite-time guarantees for nonlinear two-timescale schemes remain difficult to obtain, especially under constant step-sizes. In this paper, we study nonlinear TTSA with step-sizes α β. Under standard stability, regularity, and Markovian noise assumptions, we upper bound the mean-squared error and the bias of both iterates around their limiting equilibria. Our bounds scale as O(α+β 2/α 2), which we prove to be tight when β α 3/2. The analysis separates the contributions of initial conditions, fast-timescale tracking error, Markovian dependence, and timescale coupling, thereby clarifying the origin of the β 2/α squared term. Our results reveal qualitative differences from the linear TTSA setting previously studied, showing that nonlinear dynamics introduce additional finite-time effects that are absent in the linear case.

Related papers