Which LLM is Best for Translating Natural Language Goals to PDDL
cs.AI
Submitted: 2026-09-16
Updated: 2026-09-16
License: http://creativecommons.org/licenses/by/4.0/
The gist: Bridging the gap between human intent and machine execution remains a challenge in automated planning, where expressing goals in formal languages like PDDL restricts accessibility to non-experts.
Terminology
Abstract
Bridging the gap between human intent and machine execution remains a challenge in automated planning, where expressing goals in formal languages like PDDL restricts accessibility to non-experts. This paper empirically evaluates whether current Large Language Models (LLMs) can reliably translate natural language testing goals, written in informal language by video game testers, into well-formed PDDL targets suitable for classical planning. We present a carefully designed prompt template, integrating insights from iterative experimentation, aimed at maximizing both accuracy and response coherence from multiple state-of-the-art LLMs. Six contemporary models are systematically assessed on correctness, speed, and error tendencies using real-world, domain-specific benchmarks. All models demonstrate high correctness, exceeding 92%, with Gemini 2.5 Flash achieving the highest accuracy at 96% and the lowest incidence of false positives, while GPT-4.1 leads in response speed. Despite these advances, critical distinctions exist in model performance, and occasional failures arise from language ambiguity and limitations in domain representation. Our analysis underscores both the significant progress and ongoing gaps in enabling LLMs to act as robust bridges between natural language objectives and automated planning pipelines.
Sources
- TIC: Translate-Infer-Compile for accurate "text to plan" using LLMs and Logical Representations
- Large Language Models for Planning: A Comprehensive and Systematic Survey
- Can LLM-Reasoning Models Replace Classical Planning? A Benchmark Study
- LLMs Can't Plan, But Can Help Planning in LLM-Modulo Frameworks
- LLM+P: Empowering Large Language Models with Optimal Planning Proficiency
- Lost in the Middle: How Language Models Use Long Contexts
- Faithful Chain-of-Thought Reasoning
- Generating consistent PDDL domains with Large Language Models
- Large Language Models for Robotics: Opportunities, Challenges, and Perspectives
- Translating Natural Language to Planning Goals with Large-Language Models
- Large Language Models for Robotics: A Survey
- Planetarium: A Rigorous Benchmark for Translating Text to Structured Planning Languages
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection