DuoMind: Enabling Distributed Multi-Robot Coordination with Semantic Communication
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: I'm Rosa, and with me are Dev and Taro, guest researcher.
Dev: Today's paper: "DuoMind: Enabling Distributed Multi-Robot Coordination with Semantic Communication".
Rosa: DuoMind introduces a distributed hierarchical framework for multi-robot coordination that leverages vision-language models and vision-language-action models to enable robots to perform long-horizon tasks through semantic communication.
Dev: First, who's behind it and why it matters.
Title and authors: Rosa: So, we're looking at a paper called "DuoMind: Enabling Distributed Multi-Robot Coordination with Semantic Communication." It sounds like they’re tackling that big problem of getting multiple robots to work together on complex tasks across a distributed system.
Dev: Yeah, the title tells you exactly what it is about—coordination between robots using semantic communication. I'm curious if this framework can actually handle the real-world mess of physical coordination, or if it’s strictly for controlled lab environments.
Taro: From an autonomy standpoint, I wonder how they plan to manage the long-horizon aspects when things go wrong in a distributed setting. Does this structure hold up when one robot unexpectedly deviates from the expected plan?
Rosa: That's a fair question, Taro; I’m thinking about whether these complex coordination sequences translate well outside of a perfectly controlled lab where we can just reset everything.
Dev: Exactly, and I also have to consider the performance metrics for loop rate and latency; if the communication overhead is too high, that whole distributed system falls apart quickly.
Taro: It seems like the core idea here is decoupling reasoning from control, which should help isolate where failures happen in a complex interaction.
Rosa: Precisely, and I want to see how robust this structure actually proves itself when we push it out into more open environments than just the simulation setups they use.
Dev: And we need to keep an eye on those failure modes, especially concerning message loss or delayed semantic instructions between agents.
The paper's summary: Rosa: Looking at the summary of DuoMind, it really boils down to using a Vision-Language Model as the high-level brain that reasons about the whole task, and a Vision-Language-Action model for the low-level execution on each robot.
Dev: That separation is key, right? So, instead of one massive model trying to do everything at once, you’ve got specialized components handling different layers of complexity.
Taro: I find that decomposition interesting because it suggests that we can focus the VLM on the abstract coordination—the "what" and "when"—while the VLA handles the precise physical execution of a subtask.
Rosa: Right, and they make this coordination explicit through structured natural language messages shared between agents, which is a big step for making robot intentions clear.
Dev: Those messages sound like they are designed to be compact context packets—intention, subgoals, beliefs about the task state—which should help manage the communication load.
Taro: If we look at the results mentioned in the summary, they show effectiveness on benchmarks like RoboPoly and RoboTwin, which suggests these methods work under distributed observations.
Rosa: That’s encouraging; seeing performance metrics on established multi-robot coordination tasks is a strong indicator that this isn't just theoretical stuff.
Dev: I’m watching how they handle the uncertainty field in those messages; that part seems crucial for managing the ambiguity inherent in distributed observation.
The paper's improvements: Rosa: The authors highlight several key improvements, mainly focusing on developing RoboPoly as a benchmark specifically designed for long-horizon coordination under distributed control.
Dev: Developing a dedicated benchmark is smart; it gives us a standardized way to test if this framework actually performs well in scenarios that demand sustained cooperation rather than just quick reaction times.
Taro: I’m interested in the ablation studies they mention, specifically how removing inter-agent communication impacts the system's ability to handle task completion versus just local execution.
Rosa: That’s where it gets interesting; if removing those semantic messages leads to more conflicts and asynchronous behaviors, it proves that explicit coordination is essential for distributed systems.
Dev: And I also noted how the decoupled design allows them to integrate different action models without having to rebuild the whole coordination framework from scratch, which simplifies future upgrades.
Taro: So, if we can swap out the low-level controller for something else, as long as it follows the high-level instructions correctly, that’s a very flexible architecture for future research.
Rosa: It suggests that the hierarchical orchestration layer is more important than having a single perfect low-level controller; it provides the structure.
Conclusion: Rosa: So, to wrap up on "DuoMind: Enabling Distributed Multi-Robot Coordination with Semantic Communication," the core implication is that we can build coordinated systems by clearly separating high-level reasoning from low-level control using these VLM and VLA components.
Dev: It really shows how structured semantic communication, with fields like intention and subgoals, gives robots the necessary context to coordinate reliably across long tasks.
Taro: I think the biggest impact is on making multi-robot tasks feasible in real distributed settings because it addresses the problem of coordinating complex execution under partial observability.
Rosa: Exactly; we move past just single-robot demos into systems that can perform sustained, multi-step operations in a shared environment.
Dev: And looking ahead, the challenge will be keeping that loop rate tight enough while still processing all those semantic inputs without introducing unacceptable latency for the control actions.
Taro: I think future work should focus on how this framework handles dynamic changes in the environment where we don't have a static map or known constraints.
Hanchu Zhou, Dechen Gao, Hang Wang, Brendan Lynch, Boqi Zhao, Qiyao Ma, Raman Goyal, Junshan Zhang
University of California, Davis · Microsoft Research
cs.RO, cs.AI
Submitted: 2026-10-01
Updated: 2026-10-01
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
Importance score: 90/100
The gist: DuoMind introduces a distributed hierarchical framework for multi-robot coordination that leverages vision-language models and vision-language-action models to enable robots to perform long-horizon
Key concepts
- Hierarchical Framework
- This approach breaks down multi-robot coordination into two levels: high-level reasoning and low-level execution. A vision-language model handles the strategic planning (high level), while a vision-language action model manages precise physical movements (low level). This separation allows the system to manage complex, long tasks effectively.
- VLM Orchestrator
- The VLM acts as the high-level brain for each robot. It reasons over task instructions, local observations, and messages from other robots to decide what subtask to perform next. It generates both a low-level instruction for control and semantic messages for communication.
- Semantic Inter-Agent Communication
- Robots communicate using structured natural language instead of raw data. These messages contain four key fields: intention, subgoals, belief of the task state, and uncertainty. This allows robots to share compact contextual information about their plans and observations in a way that facilitates precise coordination.
Terminology
Summary
DuoMind introduces a distributed hierarchical framework for multi-robot coordination that leverages vision-language models and vision-language-action models to enable robots to perform long-horizon tasks through semantic communication. This approach addresses the challenge of coordinating complex, fine-grained execution across multiple agents in distributed settings by decoupling high-level reasoning from low-level control and facilitating inter-agent coordination via structured natural language messages.
How it works
DuoMind is a distributed hierarchical framework where each robot employs a VLM as an orchestrator for high-level reasoning and a VLA as an action model for low-level control. At each planning step, the orchestrator at each robot reasons over the task instruction, local observations, and messages received from other robots.
The orchestrator then generates two key outputs: a low-level instruction for the action model
and semantic messages for peer robots.
This architecture exploits the strengths of pretrained models by combining the semantic reasoning capabilities of VLMs with the precise action-generation capabilities of VLAs.
Decomposition into Hierarchical Levels
The framework decomposes multi-agent coordination into two complementary levels: high-level multi-agent reasoning and low-level execution. Each robot's VLM orchestrator determines the robot’s current subtask
and translates it into a single-agent, short-horizon instruction for low-level control.
The VLA then follows this instruction using local observations to generate fine-grained robot actions.
This division of responsibility allows the orchestrator to handle high-level understanding and inter-agent coordination,
while the action model focuses on reliable low-level control,
decoupling multi-agent reasoning from the need for the action model to possess strong reasoning capabilities.
Semantic Inter-Agent Communication
DuoMind facilitates coordination among multiple robots through structured natural language messages shared during high-level reasoning. Each shared message is structured using four fields:
-
Intention:
identical to the low-level instruction sent to the action model and describes the next subtask that the ego agent intends to execute.
-
Subgoals:
specifies the expected task state after the ego agent completes its current subtask.
-
Belief of task:
describes task-relevant information observed or inferred by the ego agent, such as the location of a key object or whether a particular subtask has been completed.
-
Uncertainty of belief: An optional field communicating
the uncertainty associated with the ego agent’s task belief.
This structured sharing provides compact contextual information about the intentions, subgoals, beliefs, and associated uncertainty of other agents,
enabling the orchestrator to generate more precise low-level instructions
and improving safety and efficiency.
Benchmark Development and Evaluation
To systematically evaluate coordination capabilities, DuoMind is paired with RoboPoly, a benchmark designed for long-horizon multi-robot coordination under distributed observations and control.
RoboPoly features seven tasks where multiple robots must coordinate and complete the task through closed-loop physical execution rather than high-level planning alone.
Experiments on RoboPoly and RoboTwin demonstrate that DuoMind improves multi-robot task performance,
with ablation studies confirming the contributions of hierarchical orchestration and semantic communication.
Ablation Studies and Model Compatibility
Extensive experiments confirm the importance of the proposed components. Ablation studies show that removing inter-agent communication leads to more frequent conflicts and asynchronous behaviors,
demonstrating that explicit message sharing is vital for coordination under partial observability. Furthermore, DuoMind's decoupled design allows for different action models to be integrated without redesigning the overall framework.
Evaluation with a weaker action model (π0) still shows that DuoMind substantially outperforms the π0-only baseline,
proving that high-level reasoning and subtask decomposition are key to improving coordination over the action-model-only setting.
The gist
DuoMind introduces a distributed hierarchical framework for multi-robot coordination that leverages vision-language models and vision-language-action models to enable robots to perform long-horizon tasks through semantic communication. This approach addresses the challenge of coordinating complex, fine-grained execution across multiple agents in distributed settings by decoupling high-level reasoning from low-level control and facilitating inter-agent coordination via structured natural language messages.
The gist
DuoMind introduces a distributed hierarchical framework for multi-robot coordination that leverages vision-language models and vision-language–action models to enable robots to perform long–horizon tasks through semantic communication. This approach addresses the challenge of coordinating complex, fine–grained execution across multiple agents in distributed settings by decoupling high-level reasoning from low–level control and facilitating interagent coordination via structured natural language messages.
The gist
DuoMind introduces a distributed hierarchical framework for multi-robot coordination that leverages vision-language models and vision-language–action models to enable robots to perform long–horizon tasks through semantic communication.
Improvements for AI systems
Here are specific improvements for AI systems based on the DuoMind framework, along with what those improved systems can achieve:
-
Improving Multi-Robot Coordination via Semantic Communication:
-
Enhancing Long-Horizon Task Execution under Distributed Control:
-
Increasing Robustness to Partial Observability and Non-Stationarity in Multi-Agent Systems:
-
Enabling Heterogeneous Robot Collaboration through Unified Language Interfaces:
-
Achieving Decoupled Reasoning and Control for Scalable Deployment:
- AI System Improvement (Multi-Robot Coordination):
The improved system, powered by DuoMind, will perform complex, long-horizon manipulation tasks requiring coordinated physical execution across multiple robots. Specifically, the system can achieve:
-
Coordinated object transfer between two or more robots in a shared workspace (e.g., passing bread between two arms during
Exchange Bread
). -
Simultaneous, collision-free manipulation of multiple objects (e.g., placing items on different plates during
Prepare Snack
) by reasoning about the peer robot's intentions before acting. -
Complex sequential operations that require handover protocols, such as one robot passing an object to another for final placement onto a target (e.g.,
Hang Bag
).
- AI System Improvement (Long-Horizon Task Execution):
The improved system will excel at tasks demanding sustained, multi-step planning where the outcome is only visible much later in time. It can achieve:
-
Successfully completing tasks that require sequential subgoals across many steps, such as preparing a meal (
Cook Pot
), where the high-level orchestrator maintains the overall task objective while generating intermediate instructions for low-level execution. -
Maintaining task coherence over extended periods by using an explicit
Belief of Task
in inter-agent messages to ensure all robots are aligned on the eventual goal, even when local observations are incomplete.
- AI System Improvement (Robustness to Partial Observability):
By utilizing structured inter-agent communication (Intention, Subgoals, Belief of Task), the system will significantly reduce coordination failures caused by distributed uncertainty. It can achieve:
-
Mitigating conflicts arising from asynchronous actions by allowing robots to infer the peer agent's immediate next action and proactively adjusting their own plans to avoid interference (as demonstrated in
Cook Pot
). -
Resilience against execution errors; if one robot deviates from the plan, its peers can use the shared
Uncertainty of Belief
signal to appropriately weigh the information and adjust their coordination strategy rather than simply failing.
- AI System Improvement (Heterogeneous Robot Collaboration):
The system will be highly adaptable to new robotic embodiments or different control policies without extensive retraining of the core coordination logic. It can achieve:
-
Seamless integration with various VLA action models (e.g., switching from policy π0.5 to π0) while maintaining high-level coordination effectiveness, proving that the hierarchical orchestration layer is more critical than the specific low-level controller.
-
Facilitating communication between robots potentially using different internal reasoning architectures by enforcing a standardized natural language interface for sharing intentions and subgoals.
- AI System Improvement (Scalability and Deployment):
The decoupled architecture allows for flexible scaling of the system's complexity without prohibitive computational overhead on the orchestrator. It can achieve:
-
Efficient deployment across diverse multi-robot scenarios, as foundation models (VLM/VLA) can be reused
off-the-shelf
with minimal adaptation, making it feasible to deploy coordinated systems in settings where heavy fine-tuning for every new task is impractical. -
Scalability across multiple agents by focusing the complexity on the hierarchical reasoning layer rather than requiring every agent to process the full joint state of all other robots simultaneously.
Abstract
Vision-language models (VLMs) and vision-language-action models (VLAs) have recently driven rapid progress in general-purpose robots, yet most progress has focused on single-robot settings. Extending these capabilities to multi-robot systems remains challenging because robots must coordinate long-horizon behaviors while maintaining reliable, fine-grained execution. We introduce DuoMind, a distributed hierarchical framework for multi-robot coordination through semantic communication. Each robot uses a VLA-based action model for low-level execution and a VLM-based orchestrator for high-level reasoning and inter-agent coordination. At each planning step, the orchestrator at each robot reasons over the task instruction, local observations, and messages received from other robots. It then generates low-level instructions for the action model and semantic messages for peer robots. This architecture exploits the complementary strengths of pretrained models by combining the semantic reasoning capabilities of VLMs with the precise action-generation capabilities of VLAs. To address the scarcity of benchmarks for multi-robot coordination, we further develop RoboPoly, a benchmark comprising long-horizon manipulation tasks that require coordinated, closed-loop execution under distributed control. Experiments on RoboPoly and RoboTwin demonstrate that DuoMind improves multi-robot task performance, while ablation studies confirm the contributions of hierarchical orchestration and semantic communication. More details are available on our project page.
Sources
- Qwen3-VL Technical Report
- RT-H: Action Hierarchies Using Language
- PaliGemma: A versatile 3B VLM for transfer
- GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
- $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
- Fast-in-Slow: A Dual-System Foundation Model Unifying Fast Manipulation within Slow Reasoning
- RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation
- MIMIC-D: Multi-modal Imitation for MultI-agent Coordination with Decentralized Diffusion Policies
- CHORUS: Decentralized Multi-Embodiment Collaboration with One VLA Policy
- ManiSkill2: A Unified Benchmark for Generalizable Manipulation Skills
- Latent Theory of Mind: A Decentralized Diffusion Architecture for Cooperative Manipulation
- A Closed-Loop Multi-Agent Framework for Robust Multi-Robot Manipulation
- $\pi^{*}_{0.6}$: a VLA That Learns From Experience
- $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
- LLM-Hanabi: Evaluating Multi-Agent Gameplays with Theory-of-Mind and Rationale Inference in Imperfect Information Collaboration Game
- Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models
- Hume: Introducing System-2 Thinking in Visual-Language-Action Model
- Leveraging Adaptive Group Negotiation for Heterogeneous Multi-Robot Collaboration with Large Language Models
- Gemini Robotics 1.5: Pushing the Frontier of Generalist Robots with Advanced Embodied Reasoning, Thinking, and Motion Transfer
- DART-LLM: Dependency-Aware Multi-Robot Task Decomposition and Execution using Large Language Models
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving