ViPlan: A Benchmark for Visual Planning with Symbolic Predicates and Vision-Language Models
cs.AI
Submitted: 2025-05-19
Updated: 2026-09-01
Comments: 9 pages, 5 figures and 1 table in the main text; 49 pages, 16 figures and 20 tables including Appendix
Journal ref: EMNLP 2026 Findings
Code: https://github.com/merlerm/ViPlan
License: http://creativecommons.org/licenses/by/4.0/
The gist: Integrating Large Language Models with symbolic planners is a promising direction for obtaining verifiable and grounded plans, with recent works extending this idea to visual domains using
Terminology
Abstract
Integrating Large Language Models with symbolic planners is a promising direction for obtaining verifiable and grounded plans, with recent works extending this idea to visual domains using Vision-Language Models (VLMs). However, an open-source benchmark for comparing these approaches under matched conditions is missing, due to a lack of visual benchmarks that support symbolic planning. We present ViPlan, the first open-source benchmark for comparing VLM-grounded symbolic approaches (VLM-as-grounder) with direct VLM planning methods (VLM-as-planner). ViPlan introduces a series of increasingly challenging tasks in two visual domains: a visual variant of the classic Blocksworld planning problem and a simulated household robotics environment. Averaged across methods, we find VLM-as-grounders to outperform direct VLM planning in Blocksworld (solving 46% of the tasks against 9%), where image grounding is both crucial and accurate. However, in the household robotics tasks, where linguistic knowledge helps, VLM-as-planner methods are greatly superior to VLM-as-grounder approaches (solving 34% of the tasks against 5%), which are hindered by partial observability. Thus, ViPlan domains capture fundamental shortcomings of both planning approaches, which we further diagnose with a qualitative failure analysis. Finally, across methods, we observe no consistent benefit from Chain-of-Thought prompting, suggesting persistent limitations in current VLMs' visual reasoning abilities.
Sources
- Do As I Can, Not As I Say: Grounding Language in Robotic Affordances
- Photo-Realistic Blocksworld Dataset
- From Pixels to Predicates: Learning Symbolic World Models via Pretrained Vision-Language Models
- DASH: Detection and Assessment of Systematic Hallucinations of VLMs
- Qwen2.5-VL Technical Report
- On the Opportunities and Risks of Foundation Models
- LLM-State: Open World State Representation for Long-horizon Task Planning with Large Language Model
- EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level Planning
- Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
- AHA: A Vision-Language-Model for Detecting and Reasoning Over Failures in Robotic Manipulation
- Understanding the planning of LLM agents: A survey
- BEHAVIOR-1K: A Human-Centered, Embodied AI Benchmark with 1,000 Everyday Activities and Realistic Simulation
- LLM+P: Empowering Large Language Models with Optimal Planning Proficiency
- A Survey on Hallucination in Large Vision-Language Models
- Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs
- VisualPredicator: Learning Abstract World Models with Neuro-Symbolic Predicates for Robot Planning
- Gemma 3 Technical Report
- VSP: Assessing the dual challenges of perception and reasoning in spatial planning tasks for VLMs
- DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding
- Grounding Classical Task Planners via Vision-Language Models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection