NovaPlan: Zero-Shot Long-Horizon Manipulation via Closed-Loop Video Language Planning

arXiv:2602.20119 · cs.RO, cs.AI, cs.CV · Submitted 2026-02-23 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: Today's paper: "NovaPlan: Zero-Shot Long-Horizon Manipulation via Closed-Loop Video Language Planning".

Dev: Given that solving long-horizon manipulation requires integrating high-level semantic reasoning with low-level physical interaction,

Rosa: First, who's behind it and why it matters.

Paper summary: Rosa: Looking at the full picture of NovaPlan: Zero-Shot Long-Horizon Manipulation via Closed-Loop Video Language Planning, the authors really drive home how integrating vision and language models with geometric grounding leads to a system that can actually execute complex, multi-step plans without prior specific training.

Dev: I think what’s most important is that it's not just about generating a video; it’s about treating the generation as an iterative part of a loop where the VLM critic constantly judges the physical consistency and semantic accuracy of the plan before any action is finalized.

Taro: The implication for future autonomy research is significant because this framework suggests that robots don't need to be pre-programmed for every possible outcome; they can reason about high-level goals and dynamically generate necessary low-level interactions on the fly, which opens up a much broader scope for deployable AI.

Rosa: It seems the title itself highlights the key contribution, emphasizing that it’s zero-shot capability achieved through that closed-loop video language planning structure.

Dev: From an engineering standpoint, I see its value in handling those failure modes we discussed earlier by having a dedicated recovery mechanism that can synthesize a local repair command when things like grasp slips happen during execution.

Taro: And for the world, the impact lies in making robots much more versatile tools; instead of being specialized for one task, they become general-purpose manipulators capable of tackling anything described in natural language.

Rosa: So we’re left with a framework that bridges the gap between abstract semantic understanding and precise physical interaction through this closed-loop process outlined in NovaPlan: Zero-Shot Long-Horizon Manipulation via Closed-Loop Video Language Planning.

Conclusion: Rosa: So we're wrapping up our discussion on NovaPlan, which is titled "Zero-Shot Long-Horizon Manipulation via Closed-Loop Video Language Planning."

Dev: Yeah, that title really sums up what the paper achieves—it tackles long sequences of actions without any prior training.

Taro: I think it really speaks to a fundamental shift in how we think about teaching robots complex tasks.

Rosa: It seems the core idea is using vision and language models together in a continuous loop to make robots do things they’ve never seen before.

Dev: From my side, the engineering challenge here is managing that closed loop effectively; the latency and how fast it can verify things are really critical for real-world use.

Taro: I'm curious about what happens when the world throws something unexpected at the robot during that long sequence.

Rosa: That's exactly where Taro wants to focus, because if it works reliably outside of a controlled lab setting, that’s where the real impact lies.

Dev: If it can handle those failure modes smoothly, like when an object slips or something gets stuck, then the reliability increases dramatically.

Taro: Exactly; we need to see how robust this system is when things misbehave in an unstructured environment.

Rosa: It makes me wonder if this kind of planning capability could eventually allow robots to handle much more complex real-world scenarios than we currently envision.

Robotics and AI Institute Carnegie Mellon University Department of Robotics and AI Carnegie Mellon University Brown University Department of Computer Science Brown University University of Pennsylvania

cs.RO, cs.AI, cs.CV

Submitted: 2026-02-23

Updated: 2026-10-07

Comments: Accepted to CoRL 2026. Project webpage: https://nova-plan.github.io/

Code: https://github.com/ModelTC/lightx2v

Project page: https://nova-plan.github.io

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 83/100

The gist: Given that solving long-horizon manipulation requires integrating high-level semantic reasoning with low-level physical interaction, NovaPlan introduces a hierarchical framework that unifies

Key concepts

Closed-Loop Video Language Planning
This is the high-level planning stage where a Vision Language Model (VLM) breaks down a complex goal into sequential language-based sub-tasks. It generates multiple visual simulations (rollouts) for each step and rigorously filters them using metrics related to target accuracy, physics adherence, motion flow, and final results. This loop allows the system to self-correct its plan before executing physical actions.
Hybrid Flow Mechanism
This mechanism adaptively switches between planning object movement and hand movement during execution. When manipulating an object, it reconstructs 3D geometry for precise object flow using depth models. If needed, it switches to hand flow when the rotation between frames is large to ensure stable control and handle situations where the target object is heavily occluded.
Verify-and-Recover Mechanism
This closes the loop by having a VLM critic evaluate every execution step. If a failure occurs, such as a grasp slip, the VLM analyzes why it failed and synthesizes an immediate corrective action. This recovery uses a fast, single-step rollout to propose local repairs that allow the robot to improvise solutions instead of relying on pre-programmed strategies.
Non-Prehensile Correction
For difficult tasks where objects get stuck, NovaPlan uses non-prehensile correction—a small external interaction like a nudge or poke. The VLM guides this by generating visual annotations and text prompts that specify precise contact points and constraints. This allows the system to synthesize a targeted video showing the correct external force needed to move the object.

Terminology

Summary

Given that solving long-horizon manipulation requires integrating high-level semantic reasoning with low-level physical interaction, NovaPlan introduces a hierarchical framework that unifies closed-loop VLM and video planning with geometrically grounded robot execution for zero-shot long-horizon manipulation.

The gist

NovaPlan, a hierarchical framework that unifies closed-loop VLM and video planning with geometrically grounded robot execution for zero-shot long-horizon manipulation.

How it works: Closed-Loop Video Language Planning

NovaPlan operates within a generate-then-verify tree search pipeline. The high-level planner uses a vision-language model (VLM) to decompose high-level instructions into a sequence of languagebased sub-tasks, and then generates multiple visual rollouts for each candidate using a video generation model. These rollouts are filtered and ranked based on four key metrics: the target metric, which verifies that the correct object is being manipulated; the physics metric, assessing adherence to physical laws like gravity; the motion metric, confirming flow direction matches language commands; and the result metric, ensuring final state alignment with sub-task outcomes. The system manages planning horizon by autonomously selecting a strategic mode (long horizon, h=N) for coupled tasks or a greedy mode (short horizon, h=1) for exploratory tasks based on semantic understanding.

How it works: Geometric Grounding and Hybrid Flow Mechanism

To translate synthesized visual plans into executable robot actions, NovaPlan employs a hybrid flow mechanism that adaptively switches between object flow and hand flow. For object flow, the system reconstructs 3D geometry using models like MoGe2 and refines it with Consistent Video Depth (CVD) optimization to ensure metric accuracy. It then extracts the object’s 6-DoF motion by solving for a rigid transform using algorithms like Kabsch. Hand flow is extracted via a hand pose estimator (HaMeR), which is further grounded using a dual-anchor calibration routine to resolve scale and projective drift artifacts. The switching logic dictates that the system switches to hand flow if the rotation magnitude between adjacent frames exceeds a threshold, ensuring stable control even under heavy occlusion.

How it works: Verification and Autonomous Recovery

The system closes the loop via a verify-and-recover mechanism. After each execution step, a VLM critic evaluates the state transition by comparing the start state, current state, and target state defined by the generated video plan. If failure is detected—such as grasp slip—the VLM analyzes the failure mode and synthesizes a corrective action. This recovery generation utilizes a single-step rollout (greedy mode) to propose a local repair command designed to transition the scene from the failed state back to the target state, allowing the system to improvise solutions that open-loop strategies cannot handle.

How it works: Non-Prehensile Correction

For low-tolerance tasks where objects get stuck, NovaPlan utilizes non-prehensile correction. When a corrective action is selected, the system spatially grounds this plan by querying the VLM to identify a single contact region for a small external interaction, such as a poke or nudge. The VLM produces both a Visual Annotation (overlaying an image with a red star at the contact point) and detailed Textual Guidance (a structured prompt P). This prompt P is highly constrained, specifying global constraints like single continuous shot, precise contact targets, and realistic hand appearance + stability constraints to guide the video generation model in synthesizing a targeted correction video.

How it works: Low-Level Action Mapping

The final step involves mapping the extracted flows into executable robot trajectories. For object flow, the recovered 6-DoF motion is transformed into end-effector poses using a grasp proposal network to establish a static transformation between the object frame and the end-effector frame. For hand flow, this trajectory is directly mapped to end-effector poses by leveraging it as a kinematic prior, ensuring stable execution even when the target object is occluded. Furthermore, for non-prehensile correction, a three-stage filtering pipeline selects optimal grasps using metrics like Contact Filtering and Collision Avoidance before ranking them based on a composite score Stotal.

How it works: Experimentation and Evaluation

NovaPlan is evaluated on three long-horizon tasks (Four-layer Block Stacking, Color Sorting, Hidden Object Search) and the Functional Manipulation Benchmark (FMB). The experiments test its Long-Horizon Robustness, Hybrid Flow Efficacy, and Zero-Shot Capability. Results demonstrate that NovaPlan can perform complex assembly tasks and exhibit dexterous error recovery behaviors without any prior demonstrations or training. For instance, in the Block Stacking task, NovaPlan switches to hand flow when needed, resulting in greater stability compared to purely object-centric methods.

Improvements for AI systems

Based on the provided scientific paper NovaPlan: Zero-Shot Long-Horizon Manipulation via Closed-Loop Video Language Planning, here are specific, high-impact improvements that can be made to AI systems, derived directly from its architecture and results.


The core improvement lies in moving robotics from brittle, pre-trained policies to robust, generalist manipulators capable of complex assembly without demonstrations.

Here are the specific enhancements:

  1. mathbfRobustness through Closed-Loop Verification and Recovery (Self-Correction):

NovaPlan's closed-loop Verify and Recover mechanism allows the robot to autonomously detect failures (e.g., grasp slippage, incorrect insertion) and generate a corrective action plan in real-time. This moves the system beyond open-loop execution where failure means task abandonment.

  1. mathbfHybrid Kinematic Prior Utilization (Object vs. Hand Flow Switching):

The dynamic switching mechanism between object flow tracking and human hand pose tracking (using HaMeR estimation) provides superior stability, especially under occlusion or geometric warping in generated videos. This allows the system to maintain stable control even when the target object is partially hidden, significantly improving execution reliability compared to purely object-centric methods.

  1. mathbfGeometric Grounding of Generated Plans (Addressing Embodiment Gap):

The introduction of a geometric calibration method (using dual-anchor routines for scale recovery and projective drift compensation) resolves the embodiment gap between generated video plans and physical robot control. This ensures that abstract, synthesized motions are mapped to physically executable end-effector trajectories, allowing complex, non-prehensile behaviors (like finger poking) to be grounded in metric reality.

  1. mathbf Hierarchical Reasoning for Long Horizons (Strategic Foresight):

The VLM planner's ability to decompose tasks into subgoals and monitor execution in a closed loop enables strategic foresight. By autonomously selecting a planning horizon (short-horizon greedy vs. long-horizon strategic mode), the system balances reactivity with the need for complex, multi-step assembly sequencing required by long tasks like FMB.

  1. mathbf Constraint-Aware Prompt Engineering for VLM Planning (Structured Reasoning):

The structured prompting methodology in Appendix A (e.g., enforcing Order of Operations tests, occlusion principles, and explicit state verification) restricts the VLM's output space to physically plausible actions. This prevents the system from proposing logically impossible or structurally unsound sequences, ensuring that the high-level plan aligns with physical reality before video generation even begins.

The improved AI system (NovaPlan 2.0) can now perform:

  1. mathbfZero-Shot Complex Assembly and Manipulation:** The robot can solve intricate, multi-step assembly tasks (like the Functional Manipulation Benchmark) that require millimeter precision and complex spatial reasoning without any prior task-specific training or demonstrations.

  2. mathbfAutonomous Error Recovery in Unstructured Environments:** When execution fails due to unexpected events (e.g., slippage, minor misalignment), the system doesn't stop; it uses a VLM critic to diagnose the failure and synthesize an immediate, localized corrective action (like a non-prehensile poke) to steer the object back on course.

  3. mathbf Robust Interaction with Novel Objects:** By leveraging hand flow as a kinematic prior, the robot can maintain stable manipulation when objects are heavily occluded or when interacting with irregular shapes unseen in its training data, effectively generalizing its interaction capabilities across diverse physical geometries.

  4. mathbf Reliable Long-Horizon Planning:** The system can handle tasks requiring long-term strategic foresight (e.g., multi-stage building) by intelligently switching between short, reactive sub-goals and long, sequential planning modes based on the task's dependency structure.

Abstract

Solving complex long-horizon robotic tasks requires joint reasoning over abstract task structure and low-level physical interaction. While combining Vision-Language Models (VLMs) and video generation models offers a promising path for zero-shot planning, their individual tendencies to hallucinate physics or violate geometric consistency often compound over time, preventing reliable real-world execution. We introduce NovaPlan, a hierarchical framework that enables robust, zero-shot long-horizon manipulation by systematically proposing, verifying, and repairing visual plans. At the high level, a VLM planner decomposes tasks and filters out dynamically inconsistent futures by verifying multiple candidate video rollouts. To translate these imagined futures into reliable physical actions, NovaPlan utilizes a hybrid geometric representation that adaptively switches between object-centric flow and human hand flow. Finally, NovaPlan closes the loop by continuously monitoring execution to verify outcomes and synthesize local, non-prehensile corrective behaviors, such as fingertip poking, when failures occur. Across diverse multi-stage tasks, NovaPlan substantially outperforms prior zero-shot systems, achieving complex assembly and dexterous error recovery entirely without task-specific training or demonstrations. Please visit our project website for additional results: https://nova-plan.github.io/

Sources

Related papers