Spatial Reasoning via Modality Switching Between Language and Symbolic Representations
cs.AI
Submitted: 2026-06-30
Updated: 2026-09-14
Comments: Accepted at EMNLP Findings 2026
Code: https://github.com/HLR/Spatial-Modality-Switching
License: http://creativecommons.org/licenses/by/4.0/
The gist: Human reasoning is inherently multimodal: when problems become difficult, we rarely think in words alone.
Terminology
Abstract
Human reasoning is inherently multimodal: when problems become difficult, we rarely think in words alone. We often externalize our reasoning by sketching diagrams or drawing grids to understand the underlying conceptual structure and avoid mistakes. Building on this premise, our research investigates: (a) whether grounding multi-hop textual-spatial stories into geometry-aware modalities, such as layouts or grids, improves reasoning compared to natural language-based inference; and (b) whether a model can decide when to rely on natural language reasoning and when to switch to a structured modality. We address these questions by introducing a switching metric based on trustworthiness and complexity signals, which estimates when grounding a spatial story into structure is likely to improve performance. This takes a first step toward principled modality selection in Large Language Model (LLM) reasoning. Across our settings, switching from natural language-based reasoning to a grid-based representation improves LLM performance by up to 42%, highlighting the importance of modality choice in shaping reasoning outcomes.
Sources
- Faithfulness vs. Plausibility: On the (Un)Reliability of Explanations from Large Language Models
- An Evaluation of ChatGPT-4's Qualitative Spatial Reasoning Capabilities in RCC-8
- Dialectical language model evaluation: An initial appraisal of the commonsense spatial reasoning abilities of LLMs
- An Empirical Study of Conformal Prediction in LLM with ASP Scaffolds for Robust Reasoning
- Route to Reason: Adaptive Routing for LLM and Reasoning Strategy Selection
- Spatial Reasoning in Multimodal Large Language Models: A Survey of Tasks, Benchmarks and Methods
- SpatialCoT: Advancing Spatial Reasoning through Coordinate Alignment and Chain-of-Thought for Embodied Task Planning
- Qwen3 Technical Report
- The Llama 3 Herd of Models
- AdaCoT: Pareto-Optimal Adaptive Chain-of-Thought Triggering via Reinforcement Learning
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection