FedGuide: Diffusion Prior Alignment and Value Baseline Guidance for Heterogeneous Federated Reinforcement Learning
cs.LG, cs.RO
Submitted: 2026-07-22
Updated: 2026-10-07
Code: https://github.com/hhhhzl/fedguide
License: http://creativecommons.org/licenses/by/4.0/
The gist: Federated Reinforcement Learning (FRL) enables collaborative policy learning across distributed agents with heterogeneous environments.
Terminology
Abstract
Federated Reinforcement Learning (FRL) enables collaborative policy learning across distributed agents with heterogeneous environments. While recent methods based on variance reduction, divergence penalization, and momentum optimization improve FRL under heterogeneous settings, they still primarily synchronize policy or value-network parameters and do not explicitly address distributional mismatch among heterogeneous clients. Therefore, we propose FedGuide, a FRL framework that uses diffusion priors as behavior models to provide personalized data supported distributions for heterogeneous local policy learning. Instead of directly averaging local policies, FedGuide aggregates those diffusion priors through Optimal-Transport Mixture-of-Experts (OT-MoE), preserving heterogeneous behavior modes in distribution space. It further develops a Distribution Correction Estimation (DICE) value baseline to provide low-variance, return-aware guidance for local policy improvement. Experiments across heterogeneous environments show that FedGuide outperforms representative FRL methods in client-average returns, final-round performance, and worst-round robustness, while maintaining stable learning under stronger heterogeneity.
Sources
- A Survey of Distributed Optimization Methods for Multi-Robot Systems
- Federated Reinforcement Learning: Techniques, Applications, and Open Challenges
- Momentum for the Win: Collaborative Federated Reinforcement Learning across Heterogeneous Environments
- FedHPD: Heterogeneous Federated Reinforcement Learning via Policy Distillation
- GenDICE: Generalized Offline Estimation of Stationary Values
- Policy-Guided Diffusion
- Prior-Guided Diffusion Planning for Offline Reinforcement Learning
- Proximal Policy Optimization Algorithms
- FedREP: A Byzantine-Robust, Communication-Efficient and Privacy-Preserving Framework for Federated Learning
- Finite-Time Analysis of On-Policy Heterogeneous Federated Reinforcement Learning
- FedMoE: Personalized Federated Learning via Heterogeneous Mixture of Experts
- HFedMoE: Resource-aware Heterogeneous Federated Learning with Mixture-of-Experts
- Heterogeneous Federated Reinforcement Learning Using Wasserstein Barycenters
- ODICE: Revealing the Mystery of Distribution Correction Estimation via Orthogonal-gradient Update
- AlgaeDICE: Policy Gradient from Arbitrary Experience
- COptiDICE: Offline Constrained Reinforcement Learning via Stationary Distribution Correction Estimation
- High-Dimensional Continuous Control Using Generalized Advantage Estimation
- OpenAI Gym
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks