DuoMind: Enabling Distributed Multi-Robot Coordination with Semantic Communication
summary
The gist
DuoMind introduces a distributed hierarchical framework for multi-robot coordination that leverages vision-language models and vision-language-action models to enable robots to perform long-horizon
In short
DuoMind is a distributed framework for multi-robot coordination using vision-language models (VLMs) and vision-language-action models (VLAs). It enables robots to complete long tasks by decoupling high-level reasoning from low-level control. Robots coordinate through structured natural language messages that share intentions, subgoals, and beliefs, improving performance in complex distributed settings.
Key concepts
- Hierarchical Framework
- This approach breaks down multi-robot coordination into two levels: high-level reasoning and low-level execution. A vision-language model handles the strategic planning (high level), while a vision-language action model manages precise physical movements (low level). This separation allows the system to manage complex, long tasks effectively.
- VLM Orchestrator
- The VLM acts as the high-level brain for each robot. It reasons over task instructions, local observations, and messages from other robots to decide what subtask to perform next. It generates both a low-level instruction for control and semantic messages for communication.
- Semantic Inter-Agent Communication
- Robots communicate using structured natural language instead of raw data. These messages contain four key fields: intention, subgoals, belief of the task state, and uncertainty. This allows robots to share compact contextual information about their plans and observations in a way that facilitates precise coordination.
Terminology used across episodes
This episode discusses
- DuoMind: Enabling Distributed Multi-Robot Coordination with Semantic Communication · Paper Radio
- Qwen3-VL Technical Report
- RT-H: Action Hierarchies Using Language
- PaliGemma: A versatile 3B VLM for transfer
- GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
- pi 0: A Vision-Language-Action Flow Model for General Robot Control
- Fast-in-Slow: A Dual-System Foundation Model Unifying Fast Manipulation within Slow Reasoning
- RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation
- MIMIC-D: Multi-modal Imitation for MultI-agent Coordination with Decentralized Diffusion Policies
- CHORUS: Decentralized Multi-Embodiment Collaboration with One VLA Policy
- ManiSkill2: A Unified Benchmark for Generalizable Manipulation Skills
- Latent Theory of Mind: A Decentralized Diffusion Architecture for Cooperative Manipulation
- A Closed-Loop Multi-Agent Framework for Robust Multi-Robot Manipulation
- pi* 0.6: a VLA That Learns From Experience
- pi 0.5: a Vision-Language-Action Model with Open-World Generalization
- LLM-Hanabi: Evaluating Multi-Agent Gameplays with Theory-of-Mind and Rationale Inference in Imperfect Information Collaboration Game
- Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models
- Hume: Introducing System-2 Thinking in Visual-Language-Action Model
- Leveraging Adaptive Group Negotiation for Heterogeneous Multi-Robot Collaboration with Large Language Models
- Gemini Robotics 1.5: Pushing the Frontier of Generalist Robots with Advanced Embodied Reasoning, Thinking, and Motion Transfer
- DART-LLM: Dependency-Aware Multi-Robot Task Decomposition and Execution using Large Language Models
The paper
DuoMind: Enabling Distributed Multi-Robot Coordination with Semantic Communication · Read on arXiv
Hanchu Zhou, Dechen Gao, Hang Wang, Brendan Lynch, Boqi Zhao, Qiyao Ma, Raman Goyal, Junshan Zhang
University of California, Davis · Microsoft Research
Vision-language models (VLMs) and vision-language-action models (VLAs) have recently driven rapid progress in general-purpose robots, yet most progress has focused on single-robot settings. Extending these capabilities to multi-robot systems remains challenging because robots must coordinate long-horizon behaviors while maintaining reliable, fine-grained execution. We introduce DuoMind, a distributed hierarchical framework for multi-robot coordination through semantic communication. Each robot uses a VLA-based action model for low-level execution and a VLM-based orchestrator for high-level reasoning and inter-agent coordination. At each planning step, the orchestrator at each robot reasons over the task instruction, local observations, and messages received from other robots. It then generates low-level instructions for the action model and semantic messages for peer robots. This architecture exploits the complementary strengths of pretrained models by combining the semantic reasoning capabilities of VLMs with the precise action-generation capabilities of VLAs. To address the scarcity of benchmarks for multi-robot coordination, we further develop RoboPoly, a benchmark comprising long-horizon manipulation tasks that require coordinated, closed-loop execution under distributed control. Experiments on RoboPoly and RoboTwin demonstrate that DuoMind improves multi-robot task performance, while ablation studies confirm the contributions of hierarchical orchestration and semantic communication. More details are available on our project page.
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: I'm Rosa, and with me are Dev and Taro, guest researcher.
Dev: Today's paper: "DuoMind: Enabling Distributed Multi-Robot Coordination with Semantic Communication".
Rosa: DuoMind introduces a distributed hierarchical framework for multi-robot coordination that leverages vision-language models and vision-language-action models to enable robots to perform long-horizon tasks through semantic communication.
Dev: First, who's behind it and why it matters.
Title and authors: Rosa: So, we're looking at a paper called "DuoMind: Enabling Distributed Multi-Robot Coordination with Semantic Communication." It sounds like they’re tackling that big problem of getting multiple robots to work together on complex tasks across a distributed system.
Dev: Yeah, the title tells you exactly what it is about—coordination between robots using semantic communication. I'm curious if this framework can actually handle the real-world mess of physical coordination, or if it’s strictly for controlled lab environments.
Taro: From an autonomy standpoint, I wonder how they plan to manage the long-horizon aspects when things go wrong in a distributed setting. Does this structure hold up when one robot unexpectedly deviates from the expected plan?
Rosa: That's a fair question, Taro; I’m thinking about whether these complex coordination sequences translate well outside of a perfectly controlled lab where we can just reset everything.
Dev: Exactly, and I also have to consider the performance metrics for loop rate and latency; if the communication overhead is too high, that whole distributed system falls apart quickly.
Taro: It seems like the core idea here is decoupling reasoning from control, which should help isolate where failures happen in a complex interaction.
Rosa: Precisely, and I want to see how robust this structure actually proves itself when we push it out into more open environments than just the simulation setups they use.
Dev: And we need to keep an eye on those failure modes, especially concerning message loss or delayed semantic instructions between agents.
The paper's summary: Rosa: Looking at the summary of DuoMind, it really boils down to using a Vision-Language Model as the high-level brain that reasons about the whole task, and a Vision-Language-Action model for the low-level execution on each robot.
Dev: That separation is key, right? So, instead of one massive model trying to do everything at once, you’ve got specialized components handling different layers of complexity.
Taro: I find that decomposition interesting because it suggests that we can focus the VLM on the abstract coordination—the "what" and "when"—while the VLA handles the precise physical execution of a subtask.
Rosa: Right, and they make this coordination explicit through structured natural language messages shared between agents, which is a big step for making robot intentions clear.
Dev: Those messages sound like they are designed to be compact context packets—intention, subgoals, beliefs about the task state—which should help manage the communication load.
Taro: If we look at the results mentioned in the summary, they show effectiveness on benchmarks like RoboPoly and RoboTwin, which suggests these methods work under distributed observations.
Rosa: That’s encouraging; seeing performance metrics on established multi-robot coordination tasks is a strong indicator that this isn't just theoretical stuff.
Dev: I’m watching how they handle the uncertainty field in those messages; that part seems crucial for managing the ambiguity inherent in distributed observation.
The paper's improvements: Rosa: The authors highlight several key improvements, mainly focusing on developing RoboPoly as a benchmark specifically designed for long-horizon coordination under distributed control.
Dev: Developing a dedicated benchmark is smart; it gives us a standardized way to test if this framework actually performs well in scenarios that demand sustained cooperation rather than just quick reaction times.
Taro: I’m interested in the ablation studies they mention, specifically how removing inter-agent communication impacts the system's ability to handle task completion versus just local execution.
Rosa: That’s where it gets interesting; if removing those semantic messages leads to more conflicts and asynchronous behaviors, it proves that explicit coordination is essential for distributed systems.
Dev: And I also noted how the decoupled design allows them to integrate different action models without having to rebuild the whole coordination framework from scratch, which simplifies future upgrades.
Taro: So, if we can swap out the low-level controller for something else, as long as it follows the high-level instructions correctly, that’s a very flexible architecture for future research.
Rosa: It suggests that the hierarchical orchestration layer is more important than having a single perfect low-level controller; it provides the structure.
Conclusion: Rosa: So, to wrap up on "DuoMind: Enabling Distributed Multi-Robot Coordination with Semantic Communication," the core implication is that we can build coordinated systems by clearly separating high-level reasoning from low-level control using these VLM and VLA components.
Dev: It really shows how structured semantic communication, with fields like intention and subgoals, gives robots the necessary context to coordinate reliably across long tasks.
Taro: I think the biggest impact is on making multi-robot tasks feasible in real distributed settings because it addresses the problem of coordinating complex execution under partial observability.
Rosa: Exactly; we move past just single-robot demos into systems that can perform sustained, multi-step operations in a shared environment.
Dev: And looking ahead, the challenge will be keeping that loop rate tight enough while still processing all those semantic inputs without introducing unacceptable latency for the control actions.
Taro: I think future work should focus on how this framework handles dynamic changes in the environment where we don't have a static map or known constraints.
More episodes
- 2610.11768-Narrow and Deep: An Ontology Tower as the Knowledge of an LLM Agent for an Industrial Equipment System
- 2610.11904-Large-Scale Partition-Based RIS Beamforming For Uplink RIS-Equipped Multi-User Systems: Asymptotic Analysis
- 2610.11885-Redefining fuel poverty: Introducing the temporal equity framework (TEF)
- 2610.11900-Reach-Stabilize Control of Control-Affine Systems with Unknown Affine Parameters
- 2610.11964-From Asymptotic to Designer-Assigned-Time Control: A Review of Stability Notions, Design Mechanisms, and Controller Architectures
- 2610.12226-Stabilization of Unidirectional First-Order PDE-ODE Coupled Systems with Boundary and Distributed Input Delays
- 2610.12028-Policy Synthesis for Finite Populations of MDP Agents under Aggregate Reach-Avoid Chance Constraints
- 2610.12103-Predefined-Time Integral Reinforcement Learning for Unknown Nonlinear Systems via Inverse-Optimal Design
- 2610.12110-Adaptive dynamic programming using Lyapunov function constraints
- 2610.12324-Convex Safety Filtering via Spectral Selection for Nonconvex Safe Sets