Gondola: Grounded Vision Language Planning for Robotic Manipulation
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: Today's paper: "Gondola: Grounded Vision Language Planning for Robotic Manipulation".
Dev: Robotic manipulation faces significant challenges in generalizing across unseen objects, environments, and tasks specified by diverse language instructions.
Rosa: First, who's behind it and why it matters.
Title and authors: Rosa: Let's start by discussing the title "Gondola: Grounded Vision Language Planning for Robotic Manipulation" and who put it out there. This title really tells you right away that this work is focused on making sure the AI's plans are tied to actual visual information, not just abstract text. Dev I think the title suggests a strong emphasis on the "grounded" aspect, which implies moving beyond purely linguistic models toward something that interacts with the physical environment.
Taro: I see it as a statement that we need to solve the problem of bridging that gap between what an LLM understands linguistically and what a robot needs to physically execute in space.
Rosa: Exactly; it's about taking those high-level instructions and making them actionable by anchoring them with visual data, which is something existing methods often struggle with because they rely too much on single-view images. Dev And the authors, Shizhe Chen, Ricardo Garcia, Paul Pacaud, Cordelia Schmid Inria from Ecole normale supérieure and CNRS are known for pushing at the intersection of vision and language research.
Taro: Their background in that intersection is relevant because they’re tackling exactly where I think we need more robust systems when dealing with complex embodied tasks.
Rosa: They introduce Gondola specifically to address the limitations of existing methods which are typically constrained by single-view input and struggle with precise object grounding, which is a major weakness in current robotic planning approaches. Dev So, the paper sets itself up as a direct response to those known shortcomings in the field.
Taro: It’s interesting that they focus on multi-view inputs as their primary mechanism to overcome those single-view constraints rather than just relying on bigger models alone.
Rosa: They build their model based on Large Language Models but fine-tuned specifically for this task, which is a smart move because it leverages the reasoning power of the LLM while tailoring its output structure to what manipulation needs. Dev That fine-tuning aspect is crucial for ensuring the output format actually works with downstream planning policies.
Taro: I’m curious about how they integrated those history plans into the model, because that suggests a more complex memory mechanism than just processing the current frame.
Rosa: They encode previously generated history plans as compact text tokens to maintain contextual awareness across sequential steps in completing manipulation tasks, which is what makes it different from simpler planning loops. Dev That’s what gives it that temporal dimension we talked about earlier, which is essential for sequencing things correctly.
Taro: So they aren't just solving the current step; they're solving the whole task sequence at once by looking backward, which is a significant methodological difference from purely feed-forward planning.
Rosa: Right, so this paper introduces a structured way to generate plans that incorporates visual grounding and temporal history into the LLM's output structure. Dev This sets the stage for us to look at exactly how they achieved this integration in the next part of the discussion.
The paper's summary: Rosa: Now let's get into what Gondola actually does, as described in the paper. Essentially, it describes a model that takes multi-view images and history plans to produce a next action plan that includes both text references and segmentation masks for the target objects and locations. Dev So, it moves past just generating abstract text descriptions by providing concrete visual targets directly in the plan format.
Taro: The paper highlights that this approach generates grounded plans with segmentation masks, which means instead of just a sequence of actions, you get explicit labels for every object and location referenced in every view.
Rosa: That’s exactly right; they show how multi-view inputs alleviate occlusions for improved three dee scene perception and how those segmentation masks offer more precise and compact grounded plans. Dev The comparison to previous work shows that captions can be less accurate, missing crucial details, which Gondola tries to avoid by using these masks.
Taro: The paper also mentions the training data construction—specifically the three datasets they created: Robot grounded planning, multi-view referring expression, and pseudo long-horizon tasks.
Rosa: That’s important because those datasets were specifically designed to train the model on procedural consistency, strengthen object grounding robustness through referring queries like “Please segment one of the object name,” and teach it to reason over compositional sequences through concatenating short task sequences. Dev So, the training wasn't just general fine-tuning; it was highly targeted data construction to build these specific capabilities.
Taro: That targeted approach makes sense; if you want generalization across novel objects or complex tasks, you need data that explicitly shows the model how those scenarios look and how they should be handled.
Rosa: Exactly; by focusing on those three distinct types of data, they are directly addressing the gaps in current models concerning procedural consistency, object grounding accuracy, and long-horizon reasoning. Dev It seems like a very deliberate effort to build a system that performs well where existing models fall short.
Taro: The goal seems to be creating a planner that can handle the complexity of real manipulation instructions by being grounded in visual reality across multiple views.
Rosa: So, Gondola is essentially an LLM-based model designed to produce plans that are both temporally aware and spatially accurate by interleaving text with precise segmentation masks. Dev It’s a complex architecture, combining the vision encoder, the LLM, and the segmentation model into one integrated system.
The paper's improvements: Rosa: Moving on to what they propose as improvements in Gondola itself, we see two main technical enhancements being suggested to enhance its capabilities. First is leveraging multi-view inputs for better spatial understanding. Dev This means the AI system can transition from coarse visual representations, like bounding boxes or points, to a higher-fidelity representation of the workspace by using all the image tokens and a specialized segmentation model.
Taro: That sounds like it allows for truly view-invariant three dee scene understanding, which is something I’ve been hoping to see more of in these generative world models.
Rosa: It suggests that by concatenating multi-view image tokens with SAM2, they can generate a precise, view-invariant three dee representation of the robot's workspace for better localization and planning even in cluttered or occluded areas. Dev That capability directly addresses the need for better spatial reasoning in a way that single-view methods just can't match.
Taro: That would mean we could actually plan around objects correctly, even if one view is partially obscured, which is a huge step toward practical application.
Rosa: Secondly, they propose incorporating history plans into the LLM's input context as compact text tokens to enable more robust, temporally aware planning for long-horizon tasks. Dev This directly addresses the need for coherent reasoning over extended sequences by keeping the past steps in mind during the decision-making process.
Taro: So this is how they are tackling distribution shift—by explicitly feeding the model its own successful history, which should make it much better at tracking progress when tasks get long.
Rosa: Right, so these two improvements focus on improving spatial awareness through multi-view input and enhancing temporal coherence by injecting historical plans into the context. Dev Those are the two main levers they pull to boost performance across different levels of generalization.
Conclusion: Rosa: So, to wrap up, we've covered how Gondola uses its multi-view images and history plans to produce grounded plans with segmentation masks, which is a significant step forward in vision-language planning. Dev It seems the overall result is a model that generates outputs that are much more precise than previous methods by including those explicit visual grounding alongside temporal context.
Taro: I think the implication here is that we might see robots handling much more complex, multi-step instructions reliably in environments where things are not perfectly set up.
Rosa: Definitely; the ability to produce those view-specific masks gives us a much clearer picture of what needs to be done physically, which helps bridge the gap between abstract language and physical execution. Dev It’s a strong demonstration that grounding plans with visual evidence makes the planning output more reliable for robotic control systems.
Taro: For me, the real impact is seeing this move toward systems that can handle dynamic, unscripted situations better than those we see today on the benchmark levels L4 tasks.
Rosa: Absolutely; Gondola represents a solid direction for future research in how we build autonomous agents that can truly generalize across different visual and task instructions. Dev It’s exciting to see a system that integrates all those components into one cohesive planning framework, even if real-world deployment still requires some refinement.
Taro: We're definitely watching this, because when these systems start handling more unpredictable situations, it will change how we think about robot autonomy.
Rosa: Well, that’s our time for this paper on Gondola: Grounded Vision Language Planning for Robotic Manipulation. Dev Thanks for tuning in; we'll be back next time.
Shizhe Chen, Ricardo Garcia, Paul Pacaud, Cordelia Schmid
inria · Ecole normale supérieure
cs.RO, cs.AI, cs.CV
Submitted: 2025-06-12
Updated: 2026-09-29
Comments: Accepted to IROS 2026
Code: https://github.com/meta-llama/llama3
Project page: https://cshizhe.github.io/projects/robot_gondola.html
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 86/100
The gist: Robotic manipulation faces significant challenges in generalizing across unseen objects, environments, and tasks specified by diverse language instructions.
Key concepts
- Grounded Vision Language Planning
- This refers to a method where an AI plan is tied directly to actual visual information. Instead of just processing language abstractly, the system anchors its plans using visual data from images, making the instructions actionable in a physical environment.
- Segmentation Masks
- Gondola produces plans that include segmentation masks for target objects and locations. This means the output is not just a sequence of actions but explicitly labels what needs to be manipulated or referenced visually in every step of the plan.
- History Plans
- The model encodes previously generated plans as compact text tokens. This mechanism provides temporal awareness, allowing the AI to maintain context and reason over long sequences by looking backward at past steps during decision-making.
- Multi-view Inputs
- Using multiple images instead of a single view helps improve spatial understanding. By combining tokens from different views with a segmentation model, the system can create a precise, view-invariant three-dimensional representation of the workspace.
Terminology
Summary
Robotic manipulation faces significant challenges in generalizing across unseen objects, environments, and tasks specified by diverse language instructions. This paper introduces Gondola, a novel grounded vision-language planning model based on Large Language Models (LLMs), designed to improve generalization capabilities in robotic manipulation. Gondola addresses the limitations of existing methods—which are often constrained by single-view input and struggle with precise object grounding—by taking multi-view images and history plans to produce the next action plan with interleaved texts and segmentation masks of target objects and locations.
Gondola Architecture
The Gondola architecture consists of three main components: an image encoder, a Large Language Model (LLM), and a segmentation model. The image encoder utilizes a pretrained Vision Transformer (ViT) InternVL-300M to generate image patch embeddings, which are then processed by a 2-layer Multi-Layer Perceptron (MLP) trained to adapt the visual features to the language space. This encoder is applied across all views, with view separation handled by a special token. The LLM adopted is InternVL-4B, fine-tuned using LoRA layers. To support grounded planning, Gondola incorporates a specialized vocabulary that includes a dedicated token to signal mask generation, along with paired delimiter tokens
and
for precise object and location references. Furthermore, previously generated history plans are encoded as compact text tokens to maintain contextual awareness across sequential steps in completing manipulation tasks.
Training Data Construction
To support the training of Gondola, three types of datasets were constructed using the RLBench simulator:
-
Robot grounded planning: This dataset is structured with semantically labeled objects and fixed procedure trajectories, resulting in ground-truth plan triplets for multi-view images. Approximately 15k training tuples are created for each GemBench training task variation.
-
Multi-view referring expression: This dataset strengthens object grounding capabilities by formulating referring queries like “Please segment one of the [object name]” with expected outputs being segmentation masks across all input view images, yielding 15k multi-view image examples and 58k referring expressions.
-
Pseudo long-horizon tasks: To enhance planning for long-horizon tasks, two short task sequences are randomly concatenated, and an LLM generates coherent joint instructions. This helps the model learn to leverage history plans to track task progress.
Training Objectives
Gondola is trained to jointly optimize plan generation and multi-view object grounding using two distinct loss functions:
- For plan generation, a cross-entropy loss is employed for next token prediction:
Lplan = −Xlog p(yi y<i, I, L, H), where yi represents tokens in the generated plan including special tokens.
- For multi-view object grounding, a joint loss of binary mask prediction and Dice loss is adopted:
Lbce = −X[Mgt(p) · log(Mpred(p)) + (1 − Mgt(p)) · log(1 − Mpred(p))]
Ldice = 1 − 2/(P p Mpred(p) · Mgt(p) + P p Mpred(p) + P p Mgt(p) + ϵ
Evaluation and Results
Gondola was evaluated on the GemBench benchmark across four levels of generalization: Level 1 (L1, new locations), Level 2 (L2, novel rigid objects), Level 3 (L3, new articulated objects), and Level 4 (L4, unseen long-horizon tasks). The evaluation includes both grounded planning assessment and task completion using a low-level motion planning policy. Comprehensive results demonstrate Gondola’s superior performance:
: Gondola outperforms the state-of-the-art LLM-based method 3D-LOTUS++ [14] by absolute 10% on average.
The model shows strong performance in grounding, achieving a mask IoU of up to 100 across various levels. Ablation studies confirm that the mask-based approach outperforms box-based variants, and multi-view inputs significantly mitigate occlusions. Furthermore, incorporating history plans boosts performance on L2 and L3 by enabling more coherent planning decisions. The pseudo long-horizon task dataset is particularly beneficial for L4, mitigating the distribution shift in history plans.
Real Robot Experiments
The model was fine-tuned on a joint dataset of RLBench and real robot demonstrations using a RealSense D435 camera setup and a UR5 robot arm. Gondola demonstrates successful predictions on seen task variations but performs worse in the real-world setting compared to simulation, primarily due to the limited amount of real robot data,
leading to frequent segmentation errors.
The results show that for fine-grained manipulation tasks, smaller action chunks yield better performance, while for long-horizon tasks benefiting from consistent planning, larger action chunks are more effective.
Improvements for AI systems
As a fastidious researcher, I have meticulously analyzed the provided paper on Gondola: Grounded Vision-Language Planning for Generalizable Robotic Manipulation.
The core innovation lies in grounding LLM plans with multi-view segmentation masks and incorporating historical plan context.
Here are the specific improvements that can be made to AI systems based on this research, categorized by technical enhancement:
)1. Multi-View Grounding and Perception Enhancement
The system can transition from coarse visual representations (bounding boxes/points) to high-fidelity, view-consistent 3D scene understanding.
The improved AI system can generate a precise, view-invariant 3D representation of the robot's workspace by leveraging the concatenation of multi-view image tokens and the specialized SAM2 segmentation model. This allows for accurate localization and manipulation planning even in cluttered or occluded environments, as evidenced by Gondola's superior Mask IoU performance (Table 1).
)2. Grounded Action Generation with Precise Object/Location Referencing
The system can move beyond generating abstract actions to producing plans that explicitly identify and segment the exact target object and destination location across all camera views.
The improved AI system can output executable plans that are not just sequences of actions, but include interleaved text references (e.g., using the special tokens like
lamp
) and corresponding segmentation masks for every required object and location in every view. This ensures the low-level motion planner receives geometrically precise targets, significantly reducing errors caused by ambiguity or partial views.
)3. Context-Aware, History-Informed Long-Horizon Reasoning
The system can maintain coherent reasoning over extended sequences of tasks by explicitly encoding past successful actions into the current planning context.
The improved AI system can incorporate a textual history of previously executed plans (H) as compact tokens into the LLM's input context. This enables it to perform more robust, temporally aware planning, particularly for long-horizon tasks (Level 4), mitigating
distribution shiftwhere models revert to purely textual plans.
)4. Robust Training Data Synthesis for Generalization
The system can be trained on diverse, synthetic data that explicitly targets the weaknesses of current models—generalization across novel objects and complex task sequences.
The improved AI system benefits from training on three specialized datasets: (1) Robot Grounded Planning (for procedural consistency), (2) Multi-View Referring Expressions (to strengthen object grounding robustness), and (3) Pseudo Long-Horizon Tasks (to teach the model to reason over compositional, multi-step sequences). This targeted data construction directly addresses generalization gaps in unseen objects, novel articulations, and long-horizon planning.
)5. Adaptive Planning Strategy Based on Task Complexity
The system can dynamically adjust its planning granularity (action chunk size) based on the immediate visual context and the predicted complexity of the next step.
The improved AI system can implement an adaptive action chunking strategy: using smaller chunks for fine-grained, reactive tasks (like contact-rich manipulation) to gain better localized perception, and larger chunks for high-level, long-horizon tasks to maintain coherent global planning. This dynamic control optimizes the trade-off between rapid correction and strategic foresight.
This improved AI system can perform:
-
Generate precise 3D scene representations using multi-view inputs for superior spatial reasoning.
-
Produce executable plans explicitly grounded with view-specific segmentation masks for objects and locations, ensuring no ambiguity in physical execution.
-
Maintain coherent, context-aware planning over very long sequences of instructions by leveraging historical plan context.
-
Achieve state-of-the-art generalization across novel object types (rigid and articulated) and complex, unseen long-horizon tasks (Level 4).
-
Execute complex manipulation tasks with higher success rates in real robots by dynamically adjusting its planning granularity based on the task's immediate demands.
Sources
- Solving Rubik's Cube with a Robot Hand
- Diffusion Policy: Visuomotor Policy Learning via Action Diffusion
- Mobile ALOHA: Learning Bimanual Mobile Manipulation with Low-Cost Whole-Body Teleoperation
- RT-1: Robotics Transformer for Real-World Control at Scale
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
- PaLM-E: An Embodied Multimodal Language Model
- Octo: An Open-Source Generalist Robot Policy
- OpenVLA: An Open-Source Vision-Language-Action Model
- AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems
- THE COLOSSEUM: A Benchmark for Evaluating Generalization for Robotic Manipulation
- VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models
- GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
- Scaling Robot Learning with Semantically Imagined Experience
- LLaRA: Supercharging Robot Learning Data for Vision-Language Policy
- RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics
- Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos
- 3D Diffuser Actor: Policy Diffusion with 3D Scene Representations
- Towards Open-World Grasping with Large Vision-Language Models
- Learning Robotic Manipulation Policies from Point Clouds with Conditional Flow Matching
- Segment Anything
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving