Neuro-Symbolic Synergy for World Modeling
cs.CL
Submitted: 2026-02-11
Updated: 2026-09-14
Comments: Camera-ready version accepted at COLM 2026
Code: https://github.com/tianyi-lab/NeSyS
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large language models (LLMs) exhibit strong general-purpose reasoning capabilities, yet they frequently hallucinate when used as world models (WMs), where strict compliance with deterministic
Terminology
Abstract
Large language models (LLMs) exhibit strong general-purpose reasoning capabilities, yet they frequently hallucinate when used as world models (WMs), where strict compliance with deterministic transition rules--particularly in corner cases--is essential. In contrast, Symbolic WMs provide logical consistency but lack semantic expressivity. To bridge this gap, we propose Neuro-Symbolic Synergy (NeSyS), a framework that integrates the probabilistic semantic priors of LLMs with executable symbolic rules to achieve both expressivity and robustness. NeSyS alternates training between the two models using trajectories inadequately explained by the other. Unlike rule-based prompting, the symbolic WM contributes candidate-level scores through log-linear reranking, without requiring the LLM to interpret rule text. Rule-guided sampling prioritizes transitions that are weakly covered by symbolic rules, using 35--60% of the training pairs while outperforming full-data supervised fine-tuning in five of six settings. Experiments on ScienceWorld, WebShop, and PlanCraft demonstrate consistent gains in WM prediction accuracy and data efficiency; one-step lookahead on open-ended WebShop also improves agent reward. Our models, rules, and code are available at https://github.com/tianyi-lab/NeSyS.
Sources
- Web Agents with World Models: Learning and Leveraging Environment Dynamics in Web Navigation
- Plancraft: an evaluation dataset for planning with LLM agents
- The Llama 3 Herd of Models
- WorldLLM: Improving LLMs' world modeling using curiosity-driven theory-making
- Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs
- Selection of LLM Fine-Tuning Data based on Orthogonal Rules
- gpt-oss-120b & gpt-oss-20b Model Card
- VisualPredicator: Learning Abstract World Models with Neuro-Symbolic Predicates for Robot Planning
- Neuro-Symbolic World Models for Adapting to Open World Novelty
- Guiding LLMs The Right Way: Fast, Non-Invasive Constrained Generation
- Neurosymbolic Grounding for Compositional World Models
- OpenAI GPT-5 System Card
- WALL-E 2.0: World Alignment by NeuroSymbolic Learning improves World Model-based LLM Agents
- ScienceWorld: Is your Agent Smarter than a 5th Grader?
- Efficient Guided Generation for Large Language Models
- Qwen2 Technical Report
- Qwen2.5 Technical Report
- LLM-Based World Models Can Make Decisions Solely, But Rigorous Evaluations are Needed
- WALL-E: World Alignment by Rule Learning Improves World Model-based LLM Agents
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering