Visuospatial Perspective Taking in Multimodal Language Models
cs.CL, cs.AI
Submitted: 2026-03-04
Updated: 2026-03-04
Journal ref: Prunty, J., Zhang, S., Quinn, P., Lian, J., Xie, X., & Cheke, L. G. (2026). Visuospatial Perspective Taking in Multimodal Language Models. Proceedings of the Annual Meeting of the Cognitive Science Society, 48
Code: https://github.com/UKGovernmentBEIS/inspect_ai
License: http://creativecommons.org/licenses/by/4.0/
The gist: As multimodal language models (MLMs) are increasingly used in social and collaborative settings, it is crucial to evaluate their perspective-taking abilities.
Terminology
Abstract
As multimodal language models (MLMs) are increasingly used in social and collaborative settings, it is crucial to evaluate their perspective-taking abilities. Existing benchmarks largely rely on text-based vignettes or static scene understanding, leaving visuospatial perspective-taking (VPT) underexplored. We adapt two evaluation tasks from human studies: the Director Task, assessing VPT in a referential communication paradigm, and the Rotating Figure Task, probing perspective-taking across angular disparities. Across tasks, MLMs show pronounced deficits in Level 2 VPT, which requires inhibiting one's own perspective to adopt another's. These results expose critical limitations in current MLMs' ability to represent and reason about alternative perspectives, with implications for their use in collaborative contexts.
Sources
- International AI Safety Report
- Embers of Autoregression: Understanding Large Language Models Through the Problem They are Trained to Solve
- Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks
- Clever Hans or Neural Theory of Mind? Stress Testing Social Reasoning in Large Language Models
- GPT-4o System Card
- The 3D-PC: a benchmark for visual perspective taking in humans and machines
- Vision Language Models See What You Want but not What You See
- Seeing Through Their Eyes: Evaluating Visual Perspective Taking in Vision Language Models
- Beyond Recognition: Evaluating Visual Perspective Taking in Vision Language Models
- Failures in Perspective-taking of Multimodal AI Systems
- Procedural Dilemma Generation for Evaluating Moral Reasoning in Humans and Language Models
- Anthropocentric bias in language model evaluation
- ReAct: Synergizing Reasoning and Acting in Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering