Robust Federated Q-Learning with Almost No Communication
cs.LG, cs.SY, eess.SY
Submitted: 2026-08-24
Updated: 2026-08-24
Comments: Accepted at the 2026 American Control Conference (ACC 2026)
Journal ref: S. Maity and A. Mitra, "Robust Federated Q-Learning with Almost No Communication," 2026 American Control Conference (ACC), New Orleans, LA, USA, 2026, pp. 462-469
License: http://creativecommons.org/licenses/by/4.0/
The gist: We consider a federated reinforcement learning setting involving M agents, all of whom interact with a common Markov Decision Process (MDP).
Terminology
Abstract
We consider a federated reinforcement learning setting involving M agents, all of whom interact with a common Markov Decision Process (MDP). The agents exchange information via a central server to learn the optimal value function. Our goal is to understand to what extent one can hope for collaborative sample-complexity speedups in such a setting, when a small fraction of the agents are adversarial and can act arbitrarily. To that end, we propose Robust Fed-Q, a federated Q-learning algorithm that blends ideas from both model-based and model-free RL, along with the median-of-means device from robust statistics. We prove that despite corruption, with high-probability, Robust Fed-Q (i) guarantees exact convergence to the optimal value function in the limit of infinite samples, and (ii) enjoys near-optimal finite-time rates that benefit from collaboration. In addition, our approach requires just (1) rounds of communication to achieve each of the above guarantees, a feature of independent interest in FL where communication is the major bottleneck.
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks