Implicit vs. Explicit Prompting Strategies for LVLMs in Referential Communication

arXiv:2606.17372 · cs.CL, cs.AI · Submitted 2026-06-16 · Read on arXiv

cs.CL, cs.AI

Submitted: 2026-06-16

Updated: 2026-09-01

Comments: 17 pages

License: http://creativecommons.org/licenses/by/4.0/

The gist: Two recent studies reach apparently contradictory conclusions about whether large vision-language models (LVLMs) can coordinate similarly to humans on efficient referring expressions.

Terminology

Abstract

Two recent studies reach apparently contradictory conclusions about whether large vision-language models (LVLMs) can coordinate similarly to humans on efficient referring expressions. We control for task differences between the studies while directly comparing their prompting styles. We replicate the finding that models can coordinate efficient referring expressions when explicitly prompted to do so, suggesting that other task differences are not responsible for divergent results. However, we also find that the same models fail to infer the need for communicative efficiency from a more implicit prompt, highlighting critical differences between how humans and AI systems communicate.

Sources

Related papers