A Decentralized Partially Observable Team Decision Methodology with Delayed Information Sharing
math.OC, cs.LG, cs.SY, eess.SY, stat.ML
Submitted: 2026-09-22
Updated: 2026-09-22
Comments: 15 pages, 2 figures
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: We study decentralized partially observable team decision problems with low-rank latent dynamics and unknown system models.
Terminology
Abstract
We study decentralized partially observable team decision problems with low-rank latent dynamics and unknown system models. The proposed framework combines team-theoretic equivalence with low-rank model representations to address cooperative decision-making in partially observable Markov decision processes without prior knowledge of the transition model. Each team member makes decisions based on local private information and delayed common information shared across the team. Using only this available information, each member learns an approximate low-rank Markov decision process and applies least-squares value iteration to compute its policy. This yields a fully decentralized learning and planning algorithm that requires neither a centralized coordinator nor centralized training. We show that the resulting member-side solutions approximate the centralized team solution: despite partial observability, unknown dynamics, and delayed common information, each member recovers the corresponding component of an approximate team-optimal policy. We further establish finite-sample performance guarantees and derive a corresponding sample-complexity bound for the proposed algorithm.
Sources
- An Introduction to Centralized Training for Decentralized Execution in Cooperative Multi-Agent Reinforcement Learning
- Value-Decomposition Networks For Cooperative Multi-Agent Learning
- Representation Learning for Online and Offline RL in Low-rank MDPs
- Spectral Representation-based Reinforcement Learning
- Representation Learning for General-sum Low-rank Markov Games
- When Does Selfishness Align with Team Goals? A Structural Analysis of Equilibrium and Optimality
Related papers
- Lions and Muons: Optimization via Stochastic Frank-Wolfe under Heavy-Tailed Noise
- Adam-HNAG: A Convergent Reformulation of Adam with Accelerated Rate
- Incremental Learning in Mirror Flows
- Online Control via Counterfactual Tracking
- Asynchronous Replanning in Two Population Linear Quadratic Mean Field Games: Information Requirements and Stability
- Petrov-Galerkin operator inference with application to stability-encouraging identification