What Matters in Orchestrating Robot Policies: A Systematic Study of Hierarchical VLA Agents
cs.RO, cs.AI, cs.LG
Submitted: 2026-06-09
Updated: 2026-09-06
License: http://creativecommons.org/licenses/by/4.0/
The gist: Hierarchical vision-language-action (Hi-VLA) systems have emerged as a promising paradigm for complex robot manipulation, by using high-level VLM planners to decompose tasks into language subgoals
Terminology
Abstract
Hierarchical vision-language-action (Hi-VLA) systems have emerged as a promising paradigm for complex robot manipulation, by using high-level VLM planners to decompose tasks into language subgoals executed by low-level VLA controllers. Despite recent empirical progress, there is a lack of unified design principles for these systems: existing Hi-VLA systems differ in how they choose and connect planners, controllers, mechanisms to switch between the two, and how observations and memory are represented in the planner. In this paper, we present a systematic study of Hi-VLA design for robot manipulation. We unify representative Hi-VLA agents under an options-style control framework and benchmark core design choices across short-horizon, long-horizon, and reasoning-intensive tasks. Our analysis distills practical principles for building Hi-VLA systems, showing how model choices and interface mechanisms jointly shape performance. Applying these principles yields a substantially stronger system than either flat VLA control or a naively designed hierarchy, across experiments both in simulation and on a real ALOHA robot. Overall, our results provide a foundation for building more capable, robust, and principled hierarchical VLA agents. More information and video at jiahenghu.github.io/hi-vla.
Sources
- Gemini Robotics: Bringing AI into the Physical World
- $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
- OpenVLA: An Open-Source Vision-Language-Action Model
- Gemini Robotics 1.5: Pushing the Frontier of Generalist Robots with Advanced Embodied Reasoning, Thinking, and Motion Transfer
- GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
- RT-H: Action Hierarchies Using Language
- $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
- Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models
- HAMSTER: Hierarchical Action Models For Open-World Robot Manipulation
- RoboOS: A Hierarchical Embodied Framework for Cross-Embodiment and Multi-Agent Collaboration
- Do As I Can, Not As I Say: Grounding Language in Robotic Affordances
- SLAC: Safe and Efficient Real-Robot Reinforcement Learning via Unsupervised Simulation Pre-Training
- HITTER: A HumanoId Table TEnnis Robot via Hierarchical Planning and Learning
- Maestro: Orchestrating Robotics Modules with Vision-Language Models for Zero-Shot Generalist Robots
- Galaxea Open-World Dataset and G0 Dual-System VLA Model
- Humanoid-VLA: Towards Universal Humanoid Control with Visual Integration
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- ${\pi}_{0.7}$: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities
- Steerable Vision-Language-Action Policies for Embodied Reasoning and Hierarchical Control
- Vision-Language Models as Success Detectors
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving