Learning to Refer from Estimated Listener Gaze
cs.CL
Submitted: 2026-09-13
Updated: 2026-09-13
Comments: Accepted at COLM 2026
Code: https://github.com/Berkeley-NLP/reg-from-gaze
License: http://creativecommons.org/licenses/by/4.0/
The gist: We propose to finetune vision-language models to generate more pragmatically optimal referring expressions by transforming observations of incremental listener comprehension, in the form of gaze
Terminology
Abstract
We propose to finetune vision-language models to generate more pragmatically optimal referring expressions by transforming observations of incremental listener comprehension, in the form of gaze scanpaths, into learning signals. During training, referring expressions are sampled from the speaker policy being optimized, conditioned on images and target referents; then, a neural listener estimating human gaze behavior maps from images and sampled referring expressions to scanpaths, each represented by a sequence of fixations, with each fixation corresponding to a word in the referring expression. We experiment with several approaches to convert fixation sequences and target referents into token- and sequence-level rewards, which are used to optimize policy parameters. Through evaluation with human listeners, we find that speaker policies trained with gaze-estimating listeners result in significantly more pragmatically-optimal references than base models, reducing sequence length from 15.4 down to 4.0 words while increasing referential success from 75.2 up to 80.0%. Our work demonstrates a promising opportunity for learning to generate utterances through language-based interaction, not only from the explicit signal of communicative success, but also from implicitly-available observations of a listener's process of comprehension.
Sources
- Qwen2.5-VL Technical Report
- PaliGemma: A versatile 3B VLM for transfer
- Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
- LoRA: Low-Rank Adaptation of Large Language Models
- Microsoft COCO: Common Objects in Context
- Improved Baselines with Visual Instruction Tuning
- Seeing Eye to AI: Human Alignment via Gaze-Based Response Rewards for Large Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering