Heterogeneous Robot Collaboration in Unstructured Environments with Grounded Generative Intelligence
cs.RO, cs.AI
Submitted: 2025-10-30
Updated: 2026-09-18
Comments: Accepted to the International Conference on Intelligent Robots and Systems (IROS) 2026
Project page: https://zacravichandran.github.io/SPINE-HT
License: http://creativecommons.org/licenses/by/4.0/
The gist: While heterogeneous teams have typically been designed for well-specified missions with known semantics, generative intelligence, i.e., large language models (LLMs) and vision language models (VLMs),
Terminology
Abstract
While heterogeneous teams have typically been designed for well-specified missions with known semantics, generative intelligence, i.e., large language models (LLMs) and vision language models (VLMs), opens the possibility of teams that infer mission-relevant semantics and subtasks given high-level natural language specifications and environmental context. However, current LLM- and VLM-enabled teaming methods typically assume well-structured and known environments, limiting performance in complex real-world settings. We address these limitations via SPINE-HT, a framework that grounds the reasoning abilities of LLMs in the evolving context of a heterogeneous robot team through a three-stage process. Given mission specifications and team capabilities in natural language, an LLM infers necessary subtasks. These subtasks are validated for feasibility and then assigned to robots based on capabilities such as traversability or perception. Finally, the framework uses feedback (semantic maps and task outcomes) acquired by the robots to update subtasks online. In simulation experiments with closed-loop perception and control, our framework achieves nearly twice the success rate compared to prior LLM-enabled heterogeneous teaming approaches. In real-world experiments with a Clearpath Jackal, a Clearpath Husky, a Boston Dynamics Spot, and a high-altitude UAV, our method achieves an 87.5% success rate in missions requiring reasoning about robot capabilities and refining subtasks with online feedback. More information is provided at: https://zacravichandran.github.io/SPINE-HT.
Sources
- DEXTER-LLM: Dynamic and Explainable Coordination of Multi-Robot Systems in Unknown Environments via Large Language Models
- Enabling Large-scale Heterogeneous Collaboration with Opportunistic Communications
- Generalized Mission Planning for Heterogeneous Multi-Robot Teams via LLM-constructed Hierarchical Trees
- Language-Grounded Hierarchical Planning and Execution with Multi-Robot 3D Scene Graphs
- Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection
- ViP-LLaVA: Making Large Multimodal Models Understand Arbitrary Visual Prompts
- ATLAS Navigator: Active Task-driven LAnguage-embedded Gaussian Splatting
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving