X-Planner: Event-Structured Task Planning for Embodied Intelligence
cs.AI
Submitted: 2026-09-21
Updated: 2026-09-21
Comments: https://github.com/X-Square-Robot/Xplanner
Code: https://github.com/X-Square-Robot/Xplanner
License: http://creativecommons.org/licenses/by/4.0/
The gist: Task planning bridges high-level instructions and executable behavior in long-horizon manipulation, yet modern Vision-Language-Action (VLA) systems often leave this intermediate structure implicit.
Terminology
Abstract
Task planning bridges high-level instructions and executable behavior in long-horizon manipulation, yet modern Vision-Language-Action (VLA) systems often leave this intermediate structure implicit. Existing chain-of-thought (CoT) planners also tend to rely on coarse task-level annotations or serialize long reasoning traces token by token. We present X-Planner, a planning front-end that addresses both the supervision and representation of embodied reasoning. Our planning data combine Ego, UMI, and teleoperation under a hierarchy granularity with source-dependent annotation depth. Takeover-time annotations and human-designed failures supervise ongoing error recognition. On the model side, a shared VLM backbone exposes two event-structured plan forms: a discrete interface that emits interpretable event states and a latent interface that relays continuous CoT states across staggered Transformer depths through Staircase Decoding. A frozen latent-to-text reconstruction objective provides a semantic anchor for the latent representation. Offline two-step planning evaluation places X-Planner second among four evaluated models on both BERTScore-F1 and a judge-based Overall score. In real-robot experiments, respectively, outperforming the evaluated baselines. These results characterize planning-text quality and downstream execution.
Sources
- Latent Reasoning VLA: Latent Thinking and Prediction for Vision-Language-Action Models
- GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
- $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
- AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems
- GR-3 Technical Report
- Training Large Language Models to Reason in a Continuous Latent Space
- FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies
- LaDi-WM: A Latent Diffusion-based World Model for Predictive Manipulation
- $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
- LaDiR: Latent Diffusion Enhances LLMs for Text Reasoning
- LaDi-RL: Latent Diffusion Reasoning Prevents Entropy Collapse in Reinforcement Learning
- DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset
- OpenVLA: An Open-Source Vision-Language-Action Model
- Causal World Modeling for Robot Control
- CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation
- LaST$_{0}$: Latent Spatio-Temporal Chain-of-Thought for Robotic Vision-Language-Action Model
- LaST-VLA: Thinking in Latent Spatio-Temporal Space for Vision-Language-Action in Autonomous Driving
- Octo: An Open-Source Generalist Robot Policy
- Qwen2.5 Technical Report
- World Action Models are Zero-shot Policies
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection