Which one is banana man? Evaluating vision-language models in multi-turn pragmatic interpretation
cs.CL
Submitted: 2026-08-30
Updated: 2026-08-30
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Flexible adaptation to context and shared pragmatic intuitions contribute to smooth human conversation.
Terminology
Abstract
Flexible adaptation to context and shared pragmatic intuitions contribute to smooth human conversation. Iterated reference games---in which players repeatedly pick out novel referents using language---present a test case for agents' ability to perform context-sensitive pragmatic reasoning in multi-turn linguistic environments. We tested humans and vision--language models on their ability to identify the intended meaning of descriptions produced in iterated reference games, varying the provided context in terms of amount, order, and relevance. While humans performed well consistently, the models we evaluated could make use of prior context to interpret humans' referring expressions, but they struggled to build up the relevant context to interpret those expressions effectively. Our results suggest that the models we evaluated lack core skills needed for efficient linguistic collaboration.
Sources
- Multi-Turn Puzzles: Evaluating Interactive Reasoning and Strategic Dialogue in LLMs
- Qwen3-VL Technical Report
- Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding
- MT-PingEval: Evaluating Multi-Turn Collaboration with Private Information Games
- Gemma 3 Technical Report
- LLMs and people both learn to form conventions -- just not with each other
- Kimi-VL Technical Report
- Olmo 3
- Beyond Language Modeling: An Exploration of Multimodal Pretraining
- Vision Language Models are Biased
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering