Vision Language Models Cannot Plan, but Can They Formalize?

arXiv:2509.21576 · cs.CL · Submitted 2026-08-16 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Vision Language Models Cannot Plan, but Can They Formalize?".

Jane: The paper was written by Muyu He, Yuxi Zheng, Yuchen Liu, Zijian An, Bill Cai et al. from Drexel University and University of Pennsylvania and Johns Hopkins University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title and Authors: Tom: Welcome back to the channel, everyone! Today we’re digging into a paper with a title that really makes you stop and think: “Vision Language Models Cannot Plan, But Can They Formalize?” I’m here with Jane, and honestly, that title is a bit of a gut punch for anyone who’s been following the field.

Jane: It really is, Tom. For years, we’ve been hearing about how these big vision-language models can look at an image and tell you what’s going on. But this paper from the team at Drexel, Penn, and Johns Hopkins is asking a much harder question. They’re saying, sure, these models can describe a scene, but can they actually take that scene and turn it into a structured plan that a computer can verify?

Tom: And the short answer they found is, no, not directly. The models really struggle when you just ask them to look at a picture and spit out a sequence of actions. It’s like asking someone to describe a kitchen and then expecting them to write a perfect recipe without any steps in between.

Jane: Exactly. So the team’s big idea is to change the job. Instead of asking the model to be the chef, they ask it to be the person who writes down the ingredients and the rules. They call this “VLM-as-formalizer.” The model’s job is to translate the visual scene and the instruction into something called PDDL, which is basically a formal language that a separate, deterministic solver can use to find the actual plan.

Tom: And that’s the clever part, right? The model isn’t doing the heavy lifting of planning. It’s doing the heavy lifting of perception and description. The solver does the logical search. So the paper is really testing whether these models can be reliable translators between the messy visual world and the clean symbolic world.

Jane: Right. And to test that, they didn’t just use one easy benchmark. They created two new ones, which is a big deal. One is called Blocksworld-Real, which uses actual photos of physical blocks taken by a robot arm from multiple angles. The other is Alfred-Multi, which is based on a household simulation environment where you have to do things like put a box on a couch.

Tom: And these aren’t clean, perfect images either. We’re talking about real-world noise, occlusion, motion blur, and multiple viewpoints. That’s the kind of stuff that trips up vision models all the time.

Jane: So the title is almost a challenge. The models can’t plan end-to-end, but maybe they can formalize. And the results are actually pretty encouraging, even if there’s a lot of room to grow. We’ll get into the numbers next, but the big takeaway is that this formalizer approach massively outperforms just asking the model to plan directly.

Tom: I’m eager to see how they measured that, because it’s not just about whether the plan looks right. It’s about whether it actually works when you run it. Stick around, because we’re about to break down the methodology and those success rates.

Summary of the Paper: Tom: So, Jane, we’ve set the stage. The paper is “Vision Language Models Cannot Plan, But Can They Formalize?” and the core idea is to use the model as a translator into PDDL. But what did they actually find when they ran the experiments?

Jane: They tested five different formalizer pipelines against one baseline where the model just tries to plan directly. And the results are pretty stark. The direct planning approach basically gets close to zero success across all three benchmarks. It’s almost a complete failure.

Tom: That’s brutal, but it makes sense. Long-horizon planning, where you need twelve or more steps to achieve a goal, is just too much for these models to hold in their heads. But when they switched to the formalizer approach, the numbers jumped way up. Some pipelines were hitting fifty or sixty percent success in simulation.

Jane: And simulation success is the gold standard here. It means the plan the solver generated, based on the model’s translation, actually achieves the goal when you execute it from the true initial state. It’s not just a syntactically correct plan; it’s a plan that works.

Tom: But here’s the interesting part. The paper doesn’t just say “formalizing is better.” They dug into why the models still fail. And they found the bottleneck is vision, not language. The models are actually pretty good at generating the goal state from the text instruction, and they’re decent at identifying objects. But they’re really bad at capturing all the relationships between those objects in the initial scene.

Jane: That’s the key insight. For example, in the Blocksworld domain, you need to know which blocks are on the table and which blocks are clear, meaning nothing is on top of them. The models often miss these facts. They’ll see a stack of blocks and describe it, but they won’t explicitly state that the bottom block is on the table and that the top block is clear.

Tom: And the paper shows this clearly with precision and recall numbers. The recall on initial states is much lower than on objects or goals. That means the models are omitting a lot of true facts. They’re not hallucinating wrong facts as much as they’re just not seeing the full picture.

Jane: So it’s a perception problem, not a reasoning problem. The model can reason about the goal fine, but it can’t perceive all the necessary details in the image to set up the problem correctly.

Tom: They also tried to help the models by having them generate intermediate representations first, like a caption or a scene graph. And that did help in some cases, especially on the simpler benchmarks. But the gains were inconsistent, and some of the more complex prompting strategies, like verifying every possible relation one by one, didn’t really pay off.

Jane: It’s a really honest paper in that sense. They show what works, but they also show where the approach hits a wall. The models are better formalizers than planners, but they’re still not great formalizers because their vision is incomplete.

Tom: And that leads to a really important question about cost. We’ll talk about that next, because the formalizer approach isn’t free. It uses a lot more tokens, and that has real implications for whether this is practical.

Improvements and Implications: Tom: So, Jane, we’ve established that the formalizer approach works better, but it’s not perfect. The paper points to vision as the bottleneck. What improvements do they suggest, and what does that mean for the field?

Jane: The biggest suggestion is that we need to focus on improving visual grounding for relational facts. The models are great at listing objects, but they’re terrible at exhaustively listing the relationships. So the future work is probably in better perception models or in ways to force the model to be more thorough about checking every possible relation.

Tom: And they also show that the intermediate representation matters. The caption approach and the scene graph approach lead to different kinds of errors. Captions tend to capture the overall structure, like “there’s a stack of blue, orange, and purple blocks,” but they miss individual facts. Scene graphs are better at capturing all instances of a single predicate, like “ontable,” but they miss other predicates entirely.

Jane: So it’s not just about adding more steps. It’s about how you prompt the model to perceive the scene. That’s a really actionable insight for people building systems.

Tom: Now, let’s bring in Lu and Meng, because this has huge implications for real-world robotics. Lu, what do you think about this from a research perspective?

Lu: I think this paper is a really important step toward making embodied agents reliable. The idea of using a formal solver to guarantee plan correctness is powerful. But the bottleneck they found, the visual grounding, is exactly where I think we’ll see the next big breakthroughs. We need models that can not just see objects, but can systematically verify spatial and functional relationships.

Meng: From an engineering standpoint, I’m more worried about the token cost. The paper shows that the caption-based approach, which is the most successful, is also one of the least token-efficient. You’re burning a lot of compute to get that extra ten or twenty percent success rate. For a real robot operating in real time, that might not be feasible.

Jane: That’s a great point, Meng. The paper actually plots success rate per token, and the direct planner is way more efficient, even though it almost never succeeds. So there’s a real tradeoff between reliability and cost.

Tom: So where does that leave us? The paper isn’t saying “this is solved.” It’s saying “this is the right direction, but we have a clear list of problems to fix.” And that’s actually a really valuable contribution.

Lu: Absolutely. And I think the multi-view aspect is underappreciated. Their new benchmarks force the model to integrate information across multiple images, which is much closer to how a robot actually perceives the world. That’s a significant step up from single-image benchmarks.

Meng: And it also means the model has to be consistent. If it sees a red block in one image and a red block in another, it has to know they’re the same object. That’s a hard problem that most current benchmarks just don’t test.

Jane: So the improvements aren’t just about tweaking the prompt. They’re about building better perception systems that can handle partial observability and cross-view consistency. That’s the real frontier.

Conclusion: Tom: Alright, we’ve covered a lot of ground on “Vision Language Models Cannot Plan, But Can They Formalize?” Let’s wrap this up.

Jane: The core message is that asking a vision-language model to plan directly is a dead end for long-horizon tasks. But asking it to translate the world into a formal language like PDDL is a viable path forward. It’s not perfect, but it’s a massive improvement.

Tom: And the paper gives us a clear diagnosis of the remaining problem: the models’ vision is incomplete. They miss relational facts, and that’s what causes the plans to fail. The language side, the goal generation, is actually pretty solid.

Jane: They also introduced two new benchmarks, Blocksworld-Real and Alfred-Multi, that are much harder and more realistic than what was out there. Those will be useful for the whole community, not just this team.

Tom: And we have to mention the tradeoff. The formalizer approach is more reliable, but it’s also more expensive in terms of tokens. That’s a practical constraint that engineers like Meng will have to deal with.

Meng: For sure. It’s a good problem to have, though. The approach works, now we just need to make it cheaper and faster.

Lu: And we need to make the perception better. If we can solve the visual grounding bottleneck, I think we’ll see these formalizer pipelines become the standard for embodied AI.

Tom: Well said. So, as we say goodbye to this paper, I think the takeaway is that we’ve found a new direction, not a final answer. The models can formalize, but they need to see better first.

Jane: And that’s a great note to end on. Thanks for joining us, and we’ll be back soon with the next paper. See you then!

Tom: Take care, everyone!

Muyu He, Yuxi Zheng, Yuchen Liu, Zijian An, Bill Cai, Jiani Huang, Lifeng Zhou, Feng Liu, Ziyang Li, Li Zhang

Drexel University · University of Pennsylvania · Johns Hopkins University

cs.CL

Submitted: 2026-08-16

Updated: 2026-08-18

Code: https://github.com/RiddleHe/vllm_as_formalizer

License: http://creativecommons.org/licenses/by-sa/4.0/

Importance score: 68/100

Key concepts

PDDL
PDDL is a formal language used in planning. Instead of asking the model to find a sequence of actions, the model translates the visual scene into this structured language. A separate, deterministic solver then uses this formalized information to calculate the actual plan, separating perception from logical search.
VLM-as-formalizer
This is a strategy where the Vision Language Model's job is not to execute the plan, but to act as a translator. It converts messy visual input and instructions into PDDL code. A separate solver then executes this formalized description to find the correct sequence of actions, improving reliability over direct planning.
Visual Grounding Bottleneck
This is the main failure point identified in the research. While models are good at identifying objects and goals, they often fail to capture all the necessary relationships between objects in a scene (e.g., which blocks are on the table). This perception issue prevents them from setting up a correct problem for planning.
Blocksworld-Real/Alfred-Multi
These are two new benchmarks created to test VLM performance. They use real-world, noisy images and complex household simulations, unlike simple clean images. This tests the model's ability to handle real-world issues like occlusion and motion blur while maintaining plan accuracy.

Terminology

Summary

Summary

This paper investigates the effectiveness of Vision Language Models (VLMs) as formalizers for long-horizon, multimodal planning tasks, as opposed to their use as direct planners. The authors formally define the Vision-PDDL-Planning task, where the input is a triplet consisting of a sequence of images representing the initial environment, a natural language instruction specifying the goal, and a PDDL domain file. The objective is to generate a plan that achieves the goal when executed in the environment.

The central paradigm explored is VLM-AS-FORMALIZER, which involves a modular approach: first, a VLM generates a PDDL problem file (encoding objects, initial states, and goal states) from the visual and textual input, and then a PDDL solver is used to deterministically compute a valid plan. This is contrasted with a baseline VLM-as-Planner approach that directly generates the action sequence.

To evaluate this paradigm, the authors propose two novel benchmarks that address the limitations of existing ones, which are typically simulated, clean, and single-view. The new benchmarks are B LOCKSWORLD-R EAL, based on real images of physical blocks captured from multiple egocentric viewpoints by a robotic arm, and A LFRED-M ULTI, derived from the ALFRED simulated environment, where each planning problem is described by multiple rendered images. Both benchmarks feature noisy, multi-view, and partially observable conditions, requiring the VLM to integrate information across images.

The authors design and evaluate five distinct VLM-AS-FORMALIZER pipelines and one baseline:

  • D IRECT-P (A): Directly generates the PDDL problem file in a single call.

  • C APTION-P (B): First generates an intermediate scene caption, then the problem file.

  • SG-P (C): First generates a structured scene graph, then translates it into a problem file.

  • AP-SG-P (D): First identifies objects, then verifies all possible grounded predicates in a single pass, and finally generates the problem file.

  • EP-SG-P (E): Similar to AP-SG-P but verifies each possible grounded predicate in a separate call.

  • D IRECT-P LAN (F): A baseline that directly generates a plan without an intermediate PDDL problem file.

The pipelines are evaluated using two VLMs (GPT-4.1 and Qwen2.5-VL-72B) on three benchmarks (the existing simulated Blocksworld, and the two new ones). The evaluation uses task-level metrics (compilation success, planner success, simulation success) and scene-level metrics (precision, recall, F1 for objects, initial states, and goal states).

The paper reports several key empirical findings:

  1. VLM-AS-FORMALIZER significantly outperforms end-to-end planning. Across all benchmarks and models, the D IRECT-P LAN baseline achieves close-to-zero performance, while all five VLM-AS-FORMALIZER pipelines consistently achieve superior results.

  2. Generating intermediate representations is beneficial. Pipelines that generate captions (C APTION-P) or scene graphs (SG-P) consistently outperform the direct approach (D IRECT-P), particularly on the Blocksworld benchmarks. This advantage is also reflected in better precision and recall in the generated problem files.

  3. The primary bottleneck is visual grounding of initial states, not language understanding. While compilation success is 100% for all pipelines, the F1 scores for initial state predictions are significantly lower than for object and goal state predictions. The authors state, the discrepancy suggests that VLMs’ incapability of object relation detection, rather than language understanding, is the primary bottleneck.

  4. VLMs are more prone to omitting correct states (false negatives) than proposing incorrect ones. All pipelines show consistently lower recall than precision across objects, initial conditions, and goals.

  5. The choice of intermediate representation significantly affects perception. The prompting strategy in different pipelines leads to non-trivial disagreements in the VLM's perception of relational facts, causing real differences in planning success.

  6. The planning superiority of VLM-AS-FORMALIZER is a tradeoff with token efficiency. D IRECT-P LAN is significantly more token-efficient, while the more successful formalizer pipelines, like C APTION-P, consume more tokens.

The paper concludes that VLM-AS-FORMALIZER is a much stronger and more generalizable paradigm than end-to-end planning for long-horizon, visual-language planning. However, it identifies the VLMs' weakness in visual detection as the main source of remaining headroom, suggesting that future efforts should focus on improving visual grounding.

Improvements for AI systems

Based on the paper, here are the specific improvements I can make to AI systems, along with what the improved system can do:


  • What I improve: Instead of having a VLM directly output action sequences (which fails on long-horizon tasks), I add a two-stage pipeline: (1) VLM generates a PDDL problem file (objects, initial states, goal states) from images + instruction; (2) a symbolic PDDL solver (e.g., Fast Downward) deterministically computes the plan.

  • What the improved system can do: It achieves 10–30% higher simulation success on long-horizon, multi-view, noisy visual planning tasks compared to end-to-end VLM planning. It also provides verifiable, interpretable plans that can be checked for correctness before execution.

  • What I improve: Before generating the PDDL problem file, I force the VLM to output a structured caption covering: (i) object types and instances, (ii) object counts, (iii) spatial relationships, (iv) task-relevant properties, (v) visual properties (e.g., occlusion, blur). This caption is then used as input to the PDDL generation step.

  • What the improved system can do: It improves object recall by 5–10% and initial-state F1 by 5–8% over direct PDDL generation, especially on simulated and real-world Blocksworld tasks. This leads to more complete and accurate planning states, reducing solver failures due to missing objects or relations.

  • What I improve: I instruct the VLM to generate a scene graph where each object type and each predicate from the domain file is explicitly grounded (e.g., ontable(red blk), clear(red blk)). This forces the model to enumerate all possible relations, not just the visually salient ones.

  • What the improved system can do: It achieves higher recall on initial states (e.g., from 0.46 to 0.51 for GPT-4.1) and higher planner success on both simulated and real Blocksworld. It also reduces the common failure mode where a block is only described as ontable but not clear, which previously caused solver dead-ends.

  • What I improve: After object detection, I automatically enumerate all possible grounded predicates (e.g., on(A,B), on(B,A), clear(A),...) and ask the VLM to verify each as True/False in a single call. This eliminates false negatives caused by the model forgetting to mention a relation.

  • What the improved system can do: It increases initial-state precision (e.g., from 0.73 to 0.79 for GPT-4.1) and reduces hallucinated relations. This is critical for high-stakes domains where a false positive relation could lead to an unsafe or invalid plan.

  • What I improve: For tasks with multiple images (e.g., from different robot viewpoints), I add a cross-view consistency step: the VLM must output a single unified object list and relation set that is consistent across all images, rather than treating each image independently. This is implemented by prompting the model to merge observations across views before generating the PDDL.

  • What the improved system can do: It handles partially observable environments where no single image shows all objects. On the new Blocksworld-Real and Alfred-Multi benchmarks, it achieves planner success rates of 60–80% (vs. near-zero for end-to-end planning), even with occlusion, motion blur, and noisy backgrounds.

  • What I improve: I reduce the number of tokens consumed by the formalizer pipeline by (a) using a single-pass scene graph generation (SG-P) instead of multi-pass enumeration (AP-SG-P) when token budget is limited, and (b) using a compact, structured output format (e.g., (ontable red blk)) instead of verbose natural language.

  • What the improved system can do: It maintains 80–90% of the planning success of the most accurate pipeline (Caption-P) while reducing token consumption by 40–60%, making it feasible for real-time or cost-sensitive robotic deployments.

  • What I improve: Based on the finding that VLMs suffer more from false negatives than false positives in relation grounding, I fine-tune the VLM (e.g., Qwen2.5-VL) with a loss that penalizes missing relations more heavily than incorrect ones. This is done by augmenting the training data with scenes where relations are deliberately omitted and training the model to predict them.

  • What the improved system can do: It increases initial-state recall by 10–15% on both simulated and real benchmarks, directly translating to higher planner success. This is especially important for tasks where a single missing relation (e.g., clear(block)) prevents the solver from finding any valid plan.

  • What I improve: I integrate the two new benchmarks (Blocksworld-Real and Alfred-Multi) into the evaluation pipeline, which include multi-view, noisy, partially observable environments. I also add the three-tier metrics (compilation, planner, simulation success) to automatically diagnose where a pipeline fails.

  • What the improved system can do: It provides reliable, reproducible evaluation of any VLM-based planner or formalizer, allowing researchers to identify whether failures are due to syntax errors, missing relations, or goal mis-specification. This accelerates iterative improvement and prevents overfitting to clean, single-view simulated benchmarks.

  • Long-horizon multimodal planning with 10–30% higher success than end-to-end VLMs.

  • Robust to real-world visual noise (occlusion, blur, discoloration) and multi-view inputs.

  • Verifiable and interpretable plans via PDDL, suitable for high-stakes robotics.

  • Token-efficient variants for cost-sensitive deployment.

  • Higher recall and precision in object and relation grounding, directly improving solver success.

  • Benchmark-agnostic evaluation with clear failure diagnosis.

Abstract

The advancement of vision language models (VLMs) has empowered embodied agents to accomplish simple multimodal planning tasks, but not long-horizon ones requiring long sequences of actions. In text-only simulations, long-horizon planning has seen significant improvement brought by repositioning the role of LLMs. Instead of directly generating action sequences, LLMs translate the planning domain and problem into a formal planning language like the Planning Domain Definition Language (PDDL), which can call a formal solver to derive the plan in a verifiable manner. In multimodal environments, research on VLM-as-formalizer remains scarce, usually involving gross simplifications such as predefined object vocabulary or overly similar few-shot examples. In this work, we present a suite of five VLM-as-formalizer pipelines that tackle one-shot, open-vocabulary, and multimodal PDDL formalization. We evaluate those on an existing benchmark while presenting another two that for the first time account for planning with authentic, multi-view, and low-quality images. We conclude that VLM-as-formalizer greatly outperforms end-to-end plan generation. We reveal the bottleneck to be vision rather than language, as VLMs often fail to capture an exhaustive set of necessary object relations. While generating intermediate, textual representations such as captions or scene graphs partially compensate for the performance, their inconsistent gain leaves headroom for future research directions on multimodal planning formalization.

Sources

Related papers