ViPlan: A Benchmark for Visual Planning with Symbolic Predicates and Vision-Language Models

arXiv:2505.13180 · cs.AI · Submitted 2025-05-19 · Read on arXiv

cs.AI

Submitted: 2025-05-19

Updated: 2026-09-01

Comments: 9 pages, 5 figures and 1 table in the main text; 49 pages, 16 figures and 20 tables including Appendix

Journal ref: EMNLP 2026 Findings

Code: https://github.com/merlerm/ViPlan

License: http://creativecommons.org/licenses/by/4.0/

The gist: Integrating Large Language Models with symbolic planners is a promising direction for obtaining verifiable and grounded plans, with recent works extending this idea to visual domains using

Terminology

Abstract

Integrating Large Language Models with symbolic planners is a promising direction for obtaining verifiable and grounded plans, with recent works extending this idea to visual domains using Vision-Language Models (VLMs). However, an open-source benchmark for comparing these approaches under matched conditions is missing, due to a lack of visual benchmarks that support symbolic planning. We present ViPlan, the first open-source benchmark for comparing VLM-grounded symbolic approaches (VLM-as-grounder) with direct VLM planning methods (VLM-as-planner). ViPlan introduces a series of increasingly challenging tasks in two visual domains: a visual variant of the classic Blocksworld planning problem and a simulated household robotics environment. Averaged across methods, we find VLM-as-grounders to outperform direct VLM planning in Blocksworld (solving 46% of the tasks against 9%), where image grounding is both crucial and accurate. However, in the household robotics tasks, where linguistic knowledge helps, VLM-as-planner methods are greatly superior to VLM-as-grounder approaches (solving 34% of the tasks against 5%), which are hindered by partial observability. Thus, ViPlan domains capture fundamental shortcomings of both planning approaches, which we further diagnose with a qualitative failure analysis. Finally, across methods, we observe no consistent benefit from Chain-of-Thought prompting, suggesting persistent limitations in current VLMs' visual reasoning abilities.

Sources

Related papers