NovaPlan: Zero-Shot Long-Horizon Manipulation via Closed-Loop Video Language Planning
summary
The gist
Given that solving long-horizon manipulation requires integrating high-level semantic reasoning with low-level physical interaction, NovaPlan introduces a hierarchical framework that unifies
In short
NovaPlan is a hierarchical framework that enables zero-shot long-horizon manipulation by combining high-level semantic reasoning with low-level physical robot execution. It uses closed-loop video planning, where a Vision Language Model decomposes tasks into sub-goals, generates visual rollouts, and verifies outcomes. The system adapts its control flow between object and hand movements to achieve complex assembly tasks autonomously.
Key concepts
- Closed-Loop Video Language Planning
- This is the high-level planning stage where a Vision Language Model (VLM) breaks down a complex goal into sequential language-based sub-tasks. It generates multiple visual simulations (rollouts) for each step and rigorously filters them using metrics related to target accuracy, physics adherence, motion flow, and final results. This loop allows the system to self-correct its plan before executing physical actions.
- Hybrid Flow Mechanism
- This mechanism adaptively switches between planning object movement and hand movement during execution. When manipulating an object, it reconstructs 3D geometry for precise object flow using depth models. If needed, it switches to hand flow when the rotation between frames is large to ensure stable control and handle situations where the target object is heavily occluded.
- Verify-and-Recover Mechanism
- This closes the loop by having a VLM critic evaluate every execution step. If a failure occurs, such as a grasp slip, the VLM analyzes why it failed and synthesizes an immediate corrective action. This recovery uses a fast, single-step rollout to propose local repairs that allow the robot to improvise solutions instead of relying on pre-programmed strategies.
- Non-Prehensile Correction
- For difficult tasks where objects get stuck, NovaPlan uses non-prehensile correction—a small external interaction like a nudge or poke. The VLM guides this by generating visual annotations and text prompts that specify precise contact points and constraints. This allows the system to synthesize a targeted video showing the correct external force needed to move the object.
Terminology used across episodes
This episode discusses
- NovaPlan: Zero-Shot Long-Horizon Manipulation via Closed-Loop Video Language Planning · Paper Radio
- Embodied Hands: Modeling and Capturing Hands and Bodies Together
- CodeDiffuser: Attention-Enhanced Diffusion Policy via VLM-Generated Code for Instruction Ambiguity
The paper
NovaPlan: Zero-Shot Long-Horizon Manipulation via Closed-Loop Video Language Planning · Read on arXiv
Robotics and AI Institute Carnegie Mellon University Department of Robotics and AI Carnegie Mellon University Brown University Department of Computer Science Brown University University of Pennsylvania
Solving complex long-horizon robotic tasks requires joint reasoning over abstract task structure and low-level physical interaction. While combining Vision-Language Models (VLMs) and video generation models offers a promising path for zero-shot planning, their individual tendencies to hallucinate physics or violate geometric consistency often compound over time, preventing reliable real-world execution. We introduce NovaPlan, a hierarchical framework that enables robust, zero-shot long-horizon manipulation by systematically proposing, verifying, and repairing visual plans. At the high level, a VLM planner decomposes tasks and filters out dynamically inconsistent futures by verifying multiple candidate video rollouts. To translate these imagined futures into reliable physical actions, NovaPlan utilizes a hybrid geometric representation that adaptively switches between object-centric flow and human hand flow. Finally, NovaPlan closes the loop by continuously monitoring execution to verify outcomes and synthesize local, non-prehensile corrective behaviors, such as fingertip poking, when failures occur. Across diverse multi-stage tasks, NovaPlan substantially outperforms prior zero-shot systems, achieving complex assembly and dexterous error recovery entirely without task-specific training or demonstrations. Please visit our project website for additional results: https://nova-plan.github.io/
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: Today's paper: "NovaPlan: Zero-Shot Long-Horizon Manipulation via Closed-Loop Video Language Planning".
Dev: Given that solving long-horizon manipulation requires integrating high-level semantic reasoning with low-level physical interaction,
Rosa: First, who's behind it and why it matters.
Paper summary: Rosa: Looking at the full picture of NovaPlan: Zero-Shot Long-Horizon Manipulation via Closed-Loop Video Language Planning, the authors really drive home how integrating vision and language models with geometric grounding leads to a system that can actually execute complex, multi-step plans without prior specific training.
Dev: I think what’s most important is that it's not just about generating a video; it’s about treating the generation as an iterative part of a loop where the VLM critic constantly judges the physical consistency and semantic accuracy of the plan before any action is finalized.
Taro: The implication for future autonomy research is significant because this framework suggests that robots don't need to be pre-programmed for every possible outcome; they can reason about high-level goals and dynamically generate necessary low-level interactions on the fly, which opens up a much broader scope for deployable AI.
Rosa: It seems the title itself highlights the key contribution, emphasizing that it’s zero-shot capability achieved through that closed-loop video language planning structure.
Dev: From an engineering standpoint, I see its value in handling those failure modes we discussed earlier by having a dedicated recovery mechanism that can synthesize a local repair command when things like grasp slips happen during execution.
Taro: And for the world, the impact lies in making robots much more versatile tools; instead of being specialized for one task, they become general-purpose manipulators capable of tackling anything described in natural language.
Rosa: So we’re left with a framework that bridges the gap between abstract semantic understanding and precise physical interaction through this closed-loop process outlined in NovaPlan: Zero-Shot Long-Horizon Manipulation via Closed-Loop Video Language Planning.
Conclusion: Rosa: So we're wrapping up our discussion on NovaPlan, which is titled "Zero-Shot Long-Horizon Manipulation via Closed-Loop Video Language Planning."
Dev: Yeah, that title really sums up what the paper achieves—it tackles long sequences of actions without any prior training.
Taro: I think it really speaks to a fundamental shift in how we think about teaching robots complex tasks.
Rosa: It seems the core idea is using vision and language models together in a continuous loop to make robots do things they’ve never seen before.
Dev: From my side, the engineering challenge here is managing that closed loop effectively; the latency and how fast it can verify things are really critical for real-world use.
Taro: I'm curious about what happens when the world throws something unexpected at the robot during that long sequence.
Rosa: That's exactly where Taro wants to focus, because if it works reliably outside of a controlled lab setting, that’s where the real impact lies.
Dev: If it can handle those failure modes smoothly, like when an object slips or something gets stuck, then the reliability increases dramatically.
Taro: Exactly; we need to see how robust this system is when things misbehave in an unstructured environment.
Rosa: It makes me wonder if this kind of planning capability could eventually allow robots to handle much more complex real-world scenarios than we currently envision.
More episodes
- 2610.12245-Fixed-Reference Pose Residuals for Measuring Cross-Dataset Cue Transfer in Human-Robot Interaction Anticipation
- 2610.12249-Real-Time Motion Planning with Dynamic Hazards: Classical vs. Learning-Based Methods
- 2610.12272-Walking on Roofs: Exploring the Potential of Walking Robots for Construction Work on Roofs
- 2610.12276-Toward Lunar Legged Robots: Field Deployment Lessons at LUNA
- 2610.12285-PLaW-VLA: Predictive Latent World Modeling for Vision-Language-Action Policies
- 2610.12368-LiteNWM: Efficient Latent World Models for Onboard Visual Navigation in the Wild
- 2610.12435-VioLA: Learning Generalist Humanoid Control Policies from Human Data
- 2610.12404-A Physics-Informed Collision Learning Framework for Collaborative Robot Motion Generation
- 2610.12411-GLIO2: A GPU-Parallelized Tightly-Coupled LiDAR-Inertial-GNSS System for Robust and Real-Time Global Localization and Mapping
- 2610.12424-RoboRSI: Stable, efficient, and reusable robot self-evolution in complex real-world environments