Vision Language Models Cannot Plan, but Can They Formalize?
summary
In short
This episode discusses a paper exploring whether Vision Language Models can create structured plans for complex tasks. The hosts conclude that while direct planning fails, using a 'formalizer' approach—translating visual scenes into PDDL—is highly reliable. However, the models' incomplete vision remains the primary bottleneck.
Key concepts
- PDDL
- PDDL is a formal language used in planning. Instead of asking the model to find a sequence of actions, the model translates the visual scene into this structured language. A separate, deterministic solver then uses this formalized information to calculate the actual plan, separating perception from logical search.
- VLM-as-formalizer
- This is a strategy where the Vision Language Model's job is not to execute the plan, but to act as a translator. It converts messy visual input and instructions into PDDL code. A separate solver then executes this formalized description to find the correct sequence of actions, improving reliability over direct planning.
- Visual Grounding Bottleneck
- This is the main failure point identified in the research. While models are good at identifying objects and goals, they often fail to capture all the necessary relationships between objects in a scene (e.g., which blocks are on the table). This perception issue prevents them from setting up a correct problem for planning.
- Blocksworld-Real/Alfred-Multi
- These are two new benchmarks created to test VLM performance. They use real-world, noisy images and complex household simulations, unlike simple clean images. This tests the model's ability to handle real-world issues like occlusion and motion blur while maintaining plan accuracy.
Terminology used across episodes
This episode discusses
- Vision Language Models Cannot Plan, but Can They Formalize? · Paper Radio
- Do As I Can, Not As I Say: Grounding Language in Robotic Affordances
- Domain-Conditioned Scene Graphs for State-Grounded Task Planning
- Look Before You Leap: Unveiling the Power of GPT-4V in Robotic Vision-Language Planning
- VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models
- FEAST: A Flexible Mealtime-Assistance System Towards In-the-Wild Personalization
- Unifying Inference-Time Planning Language Generation
- DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset
- Graph Density-Aware Losses for Novel Compositions in Scene Graph Generation
- AI2-THOR: An Interactive 3D Environment for Visual AI
- Bilevel Learning for Bilevel Planning
- VisualAgentBench: Towards Large Multimodal Models as Visual Foundation Agents
- ViPlan: A Benchmark for Visual Planning with Symbolic Predicates and Vision-Language Models · Paper Radio
- Vision-Language Interpreter for Robot Task Planning
- ALFWorld: Aligning Text and Embodied Environments for Interactive Learning
- Grounded Vision-Language Interpreter for Long-Horizon Bimanual Task and Motion Planning
- Predicate Debiasing in Vision-Language Models Integration for Scene Graph Generation Enhancement
- CHALET: Cornell House Agent Learning Environment
- EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents
- LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks
- VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks
The paper
Vision Language Models Cannot Plan, but Can They Formalize? · Read on arXiv
Muyu He, Yuxi Zheng, Yuchen Liu, Zijian An, Bill Cai, Jiani Huang, Lifeng Zhou, Feng Liu, Ziyang Li, Li Zhang
Drexel University · University of Pennsylvania · Johns Hopkins University
The advancement of vision language models (VLMs) has empowered embodied agents to accomplish simple multimodal planning tasks, but not long-horizon ones requiring long sequences of actions. In text-only simulations, long-horizon planning has seen significant improvement brought by repositioning the role of LLMs. Instead of directly generating action sequences, LLMs translate the planning domain and problem into a formal planning language like the Planning Domain Definition Language (PDDL), which can call a formal solver to derive the plan in a verifiable manner. In multimodal environments, research on VLM-as-formalizer remains scarce, usually involving gross simplifications such as predefined object vocabulary or overly similar few-shot examples. In this work, we present a suite of five VLM-as-formalizer pipelines that tackle one-shot, open-vocabulary, and multimodal PDDL formalization. We evaluate those on an existing benchmark while presenting another two that for the first time account for planning with authentic, multi-view, and low-quality images. We conclude that VLM-as-formalizer greatly outperforms end-to-end plan generation. We reveal the bottleneck to be vision rather than language, as VLMs often fail to capture an exhaustive set of necessary object relations. While generating intermediate, textual representations such as captions or scene graphs partially compensate for the performance, their inconsistent gain leaves headroom for future research directions on multimodal planning formalization.
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Vision Language Models Cannot Plan, but Can They Formalize?".
Jane: The paper was written by Muyu He, Yuxi Zheng, Yuchen Liu, Zijian An, Bill Cai et al. from Drexel University and University of Pennsylvania and Johns Hopkins University.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title and Authors: Tom: Welcome back to the channel, everyone! Today we’re digging into a paper with a title that really makes you stop and think: “Vision Language Models Cannot Plan, But Can They Formalize?” I’m here with Jane, and honestly, that title is a bit of a gut punch for anyone who’s been following the field.
Jane: It really is, Tom. For years, we’ve been hearing about how these big vision-language models can look at an image and tell you what’s going on. But this paper from the team at Drexel, Penn, and Johns Hopkins is asking a much harder question. They’re saying, sure, these models can describe a scene, but can they actually take that scene and turn it into a structured plan that a computer can verify?
Tom: And the short answer they found is, no, not directly. The models really struggle when you just ask them to look at a picture and spit out a sequence of actions. It’s like asking someone to describe a kitchen and then expecting them to write a perfect recipe without any steps in between.
Jane: Exactly. So the team’s big idea is to change the job. Instead of asking the model to be the chef, they ask it to be the person who writes down the ingredients and the rules. They call this “VLM-as-formalizer.” The model’s job is to translate the visual scene and the instruction into something called PDDL, which is basically a formal language that a separate, deterministic solver can use to find the actual plan.
Tom: And that’s the clever part, right? The model isn’t doing the heavy lifting of planning. It’s doing the heavy lifting of perception and description. The solver does the logical search. So the paper is really testing whether these models can be reliable translators between the messy visual world and the clean symbolic world.
Jane: Right. And to test that, they didn’t just use one easy benchmark. They created two new ones, which is a big deal. One is called Blocksworld-Real, which uses actual photos of physical blocks taken by a robot arm from multiple angles. The other is Alfred-Multi, which is based on a household simulation environment where you have to do things like put a box on a couch.
Tom: And these aren’t clean, perfect images either. We’re talking about real-world noise, occlusion, motion blur, and multiple viewpoints. That’s the kind of stuff that trips up vision models all the time.
Jane: So the title is almost a challenge. The models can’t plan end-to-end, but maybe they can formalize. And the results are actually pretty encouraging, even if there’s a lot of room to grow. We’ll get into the numbers next, but the big takeaway is that this formalizer approach massively outperforms just asking the model to plan directly.
Tom: I’m eager to see how they measured that, because it’s not just about whether the plan looks right. It’s about whether it actually works when you run it. Stick around, because we’re about to break down the methodology and those success rates.
Summary of the Paper: Tom: So, Jane, we’ve set the stage. The paper is “Vision Language Models Cannot Plan, But Can They Formalize?” and the core idea is to use the model as a translator into PDDL. But what did they actually find when they ran the experiments?
Jane: They tested five different formalizer pipelines against one baseline where the model just tries to plan directly. And the results are pretty stark. The direct planning approach basically gets close to zero success across all three benchmarks. It’s almost a complete failure.
Tom: That’s brutal, but it makes sense. Long-horizon planning, where you need twelve or more steps to achieve a goal, is just too much for these models to hold in their heads. But when they switched to the formalizer approach, the numbers jumped way up. Some pipelines were hitting fifty or sixty percent success in simulation.
Jane: And simulation success is the gold standard here. It means the plan the solver generated, based on the model’s translation, actually achieves the goal when you execute it from the true initial state. It’s not just a syntactically correct plan; it’s a plan that works.
Tom: But here’s the interesting part. The paper doesn’t just say “formalizing is better.” They dug into why the models still fail. And they found the bottleneck is vision, not language. The models are actually pretty good at generating the goal state from the text instruction, and they’re decent at identifying objects. But they’re really bad at capturing all the relationships between those objects in the initial scene.
Jane: That’s the key insight. For example, in the Blocksworld domain, you need to know which blocks are on the table and which blocks are clear, meaning nothing is on top of them. The models often miss these facts. They’ll see a stack of blocks and describe it, but they won’t explicitly state that the bottom block is on the table and that the top block is clear.
Tom: And the paper shows this clearly with precision and recall numbers. The recall on initial states is much lower than on objects or goals. That means the models are omitting a lot of true facts. They’re not hallucinating wrong facts as much as they’re just not seeing the full picture.
Jane: So it’s a perception problem, not a reasoning problem. The model can reason about the goal fine, but it can’t perceive all the necessary details in the image to set up the problem correctly.
Tom: They also tried to help the models by having them generate intermediate representations first, like a caption or a scene graph. And that did help in some cases, especially on the simpler benchmarks. But the gains were inconsistent, and some of the more complex prompting strategies, like verifying every possible relation one by one, didn’t really pay off.
Jane: It’s a really honest paper in that sense. They show what works, but they also show where the approach hits a wall. The models are better formalizers than planners, but they’re still not great formalizers because their vision is incomplete.
Tom: And that leads to a really important question about cost. We’ll talk about that next, because the formalizer approach isn’t free. It uses a lot more tokens, and that has real implications for whether this is practical.
Improvements and Implications: Tom: So, Jane, we’ve established that the formalizer approach works better, but it’s not perfect. The paper points to vision as the bottleneck. What improvements do they suggest, and what does that mean for the field?
Jane: The biggest suggestion is that we need to focus on improving visual grounding for relational facts. The models are great at listing objects, but they’re terrible at exhaustively listing the relationships. So the future work is probably in better perception models or in ways to force the model to be more thorough about checking every possible relation.
Tom: And they also show that the intermediate representation matters. The caption approach and the scene graph approach lead to different kinds of errors. Captions tend to capture the overall structure, like “there’s a stack of blue, orange, and purple blocks,” but they miss individual facts. Scene graphs are better at capturing all instances of a single predicate, like “ontable,” but they miss other predicates entirely.
Jane: So it’s not just about adding more steps. It’s about how you prompt the model to perceive the scene. That’s a really actionable insight for people building systems.
Tom: Now, let’s bring in Lu and Meng, because this has huge implications for real-world robotics. Lu, what do you think about this from a research perspective?
Lu: I think this paper is a really important step toward making embodied agents reliable. The idea of using a formal solver to guarantee plan correctness is powerful. But the bottleneck they found, the visual grounding, is exactly where I think we’ll see the next big breakthroughs. We need models that can not just see objects, but can systematically verify spatial and functional relationships.
Meng: From an engineering standpoint, I’m more worried about the token cost. The paper shows that the caption-based approach, which is the most successful, is also one of the least token-efficient. You’re burning a lot of compute to get that extra ten or twenty percent success rate. For a real robot operating in real time, that might not be feasible.
Jane: That’s a great point, Meng. The paper actually plots success rate per token, and the direct planner is way more efficient, even though it almost never succeeds. So there’s a real tradeoff between reliability and cost.
Tom: So where does that leave us? The paper isn’t saying “this is solved.” It’s saying “this is the right direction, but we have a clear list of problems to fix.” And that’s actually a really valuable contribution.
Lu: Absolutely. And I think the multi-view aspect is underappreciated. Their new benchmarks force the model to integrate information across multiple images, which is much closer to how a robot actually perceives the world. That’s a significant step up from single-image benchmarks.
Meng: And it also means the model has to be consistent. If it sees a red block in one image and a red block in another, it has to know they’re the same object. That’s a hard problem that most current benchmarks just don’t test.
Jane: So the improvements aren’t just about tweaking the prompt. They’re about building better perception systems that can handle partial observability and cross-view consistency. That’s the real frontier.
Conclusion: Tom: Alright, we’ve covered a lot of ground on “Vision Language Models Cannot Plan, But Can They Formalize?” Let’s wrap this up.
Jane: The core message is that asking a vision-language model to plan directly is a dead end for long-horizon tasks. But asking it to translate the world into a formal language like PDDL is a viable path forward. It’s not perfect, but it’s a massive improvement.
Tom: And the paper gives us a clear diagnosis of the remaining problem: the models’ vision is incomplete. They miss relational facts, and that’s what causes the plans to fail. The language side, the goal generation, is actually pretty solid.
Jane: They also introduced two new benchmarks, Blocksworld-Real and Alfred-Multi, that are much harder and more realistic than what was out there. Those will be useful for the whole community, not just this team.
Tom: And we have to mention the tradeoff. The formalizer approach is more reliable, but it’s also more expensive in terms of tokens. That’s a practical constraint that engineers like Meng will have to deal with.
Meng: For sure. It’s a good problem to have, though. The approach works, now we just need to make it cheaper and faster.
Lu: And we need to make the perception better. If we can solve the visual grounding bottleneck, I think we’ll see these formalizer pipelines become the standard for embodied AI.
Tom: Well said. So, as we say goodbye to this paper, I think the takeaway is that we’ve found a new direction, not a final answer. The models can formalize, but they need to see better first.
Jane: And that’s a great note to end on. Thanks for joining us, and we’ll be back soon with the next paper. See you then!
Tom: Take care, everyone!
More episodes
- 2610.10857-Self-Supervised Keyframe Discovery for Horizon-Invariant Behavior Cloning
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization