Implicit vs. Explicit Prompting Strategies for LVLMs in Referential Communication
cs.CL, cs.AI
Submitted: 2026-06-16
Updated: 2026-09-01
Comments: 17 pages
License: http://creativecommons.org/licenses/by/4.0/
The gist: Two recent studies reach apparently contradictory conclusions about whether large vision-language models (LVLMs) can coordinate similarly to humans on efficient referring expressions.
Terminology
Abstract
Two recent studies reach apparently contradictory conclusions about whether large vision-language models (LVLMs) can coordinate similarly to humans on efficient referring expressions. We control for task differences between the studies while directly comparing their prompting styles. We replicate the finding that models can coordinate efficient referring expressions when explicitly prompted to do so, suggesting that other task differences are not responsible for divergent results. However, we also find that the same models fail to infer the need for communicative efficiency from a more implicit prompt, highlighting critical differences between how humans and AI systems communicate.
Sources
- Post-training for Efficient Communication via Convention Formation
- LLMs and people both learn to form conventions -- just not with each other
- A Benchmark to Assess Common Ground in Human-AI Collaboration
- Context informs pragmatic interpretation in vision-language models
- LVLMs and Humans Ground Differently in Referential Communication
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering