An Initial Introduction to Cooperative Multi-Agent Reinforcement Learning
cs.LG, cs.MA
Submitted: 2024-05-10
Updated: 2026-09-24
License: http://creativecommons.org/licenses/by/4.0/
The gist: Multi-agent reinforcement learning (MARL) has exploded in popularity in recent years.
Terminology
Abstract
Multi-agent reinforcement learning (MARL) has exploded in popularity in recent years. While numerous approaches have been developed, they can be broadly categorized into three main types: centralized training and execution (CTE), centralized training for decentralized execution (CTDE), and decentralized training and execution (DTE). CTE methods assume centralization during training and execution (e.g., with fast, free, and perfect communication) and have the most information during execution. CTDE methods are the most common, as they leverage centralized information during training while enabling decentralized execution -- using only information available to that agent during execution. Decentralized training and execution methods make the fewest assumptions and are often simple to implement. This text is an introduction to cooperative MARL -- MARL in which all agents share a single, joint reward. It is meant to explain the setting, basic concepts, and common methods for the CTE, CTDE, and DTE settings. It does not cover all work in cooperative MARL as the area is quite extensive. I have included work that I believe is important for understanding the main concepts in the area and apologize to those that I have omitted. Topics include simple applications of single-agent methods to CTE as well as some more scalable methods that exploit the multi-agent structure, independent Q-learning and policy gradient methods and their extensions, as well as value function factorization methods including the well-known VDN, QMIX, and QPLEX approaches, and centralized critic methods including MADDPG, COMA, and MAPPO. I also discuss common misconceptions, the relationship between different approaches, and some open questions.
Sources
- Is Independent Learning All You Need in the StarCraft Multi-Agent Challenge?
- Deep Transformer Q-Networks for Partially Observable Reinforcement Learning
- Deep Recurrent Q-Learning for Partially Observable MDPs
- Best Possible Q-Learning
- Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems
- On Centralized Critics in Multi-Agent Reinforcement Learning
- The StarCraft Multi-Agent Challenge
- Proximal Policy Optimization Algorithms
- Value-Decomposition Networks For Cooperative Multi-Agent Learning
- Revisiting Parameter Sharing in Multi-Agent Deep Reinforcement Learning
- Model-based Multi-agent Reinforcement Learning: Recent Progress and Prospects
- Qatten: A General Framework for Cooperative Multiagent Reinforcement Learning
- Multi-Agent Reinforcement Learning for Autonomous Driving: A Survey
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks