Robust Federated Q-Learning with Almost No Communication

arXiv:2609.20174 · cs.LG, cs.SY, eess.SY · Submitted 2026-08-24 · Read on arXiv

cs.LG, cs.SY, eess.SY

Submitted: 2026-08-24

Updated: 2026-08-24

Comments: Accepted at the 2026 American Control Conference (ACC 2026)

Journal ref: S. Maity and A. Mitra, "Robust Federated Q-Learning with Almost No Communication," 2026 American Control Conference (ACC), New Orleans, LA, USA, 2026, pp. 462-469

License: http://creativecommons.org/licenses/by/4.0/

The gist: We consider a federated reinforcement learning setting involving M agents, all of whom interact with a common Markov Decision Process (MDP).

Terminology

Abstract

We consider a federated reinforcement learning setting involving M agents, all of whom interact with a common Markov Decision Process (MDP). The agents exchange information via a central server to learn the optimal value function. Our goal is to understand to what extent one can hope for collaborative sample-complexity speedups in such a setting, when a small fraction of the agents are adversarial and can act arbitrarily. To that end, we propose Robust Fed-Q, a federated Q-learning algorithm that blends ideas from both model-based and model-free RL, along with the median-of-means device from robust statistics. We prove that despite corruption, with high-probability, Robust Fed-Q (i) guarantees exact convergence to the optimal value function in the limit of infinite samples, and (ii) enjoys near-optimal finite-time rates that benefit from collaboration. In addition, our approach requires just (1) rounds of communication to achieve each of the above guarantees, a feature of independent interest in FL where communication is the major bottleneck.

Related papers