Gaze Target Estimation Anywhere with Concepts
Xu Cao, Houze Yang, Vipin Gunda, Zhongyi Zhou, Tianyu Xu, Adarsh Kowdle, Inki Kim, James M. Rehg
University of Illinois Urbana-Champaign · Google
cs.CV, cs.AI
Submitted: 2026-08-11
Updated: 2026-08-13
Comments: CVPR 2026 Code and Benchmark are aviliable at https://github.com/IrohXu/GazeAnywhere and https://huggingface.co/datasets/IrohXu/Gaze-Co-Benchmark
Code: https://github.com/IrohXu/GazeAnywhere
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
Importance score: 75/100
The gist: This paper introduces the Promptable Gaze Target Estimation (PGE) task and the GazeAnywhere model, a new end-to-end, concept-driven paradigm for gaze analysis.
Terminology
Summary
This paper introduces the Promptable Gaze Target Estimation (PGE) task and the GazeAnywhere model, a new end-to-end, concept-driven paradigm for gaze analysis. The authors state: "We introduce the Promptable Gaze Target Estimation (PGE) task, a new end-to-end, concept-driven paradigm for gaze analysis. PGE conditions gaze prediction on flexible user text or visual prompts (e.g., 'the boy in the red shirt' or 'person in point [0.52, 0.48]') to identify a specific subject for gaze analysis."
The paper's main contributions are summarized as: "(1) We define the Promptable Gaze Target Estimation (PGE) task, which extends the traditional gaze target estimation problem to an unconstrained, text promptable end-to-end paradigm. (2) We design a scalable data engine to generate 120K high quality PGE training annotations consisting of subject text to gaze alignment data pairs. (3) We introduce GazeAnywhere, the first promptable concept-driven gaze target estimation model."
The authors motivate the work by noting that "Existing approaches primarily employ brittle, multi-stage pipelines that require explicit inputs, like head bounding boxes and human pose, in order to identify the subject of gaze analysis. As a result, detection errors can cascade and lead to failure. Moreover, these prior works lack the flexibility of specifying the gaze analysis task via natural language prompting."
The GazeAnywhere architecture is described as follows: "The model consists of a frozen image encoder, a frozen text encoder to proceed visual modality and text modality. A transformer-based detector is used to learn joint representations and map text prompt into the main gaze target estimation and auxiliary tasks. Specifically, it uses
a frozen ViT, denoted ϕV (·), as our image encoder and
a frozen text encoder, ϕT (·), which consists of a series of transformer blocks and a final linear layer. The model includes
a head token for
explicitly predict the head localization of the prompted subject and a
Target Presence Token for
the in/out-of-frame gaze target boolean prediction objective. The detector transformer
fuse[s] these representations and refine[s] them for the gaze target estimation task and
outputs a refined sequence of tokens which are passed to
three distinct prediction heads": a Gaze Tracker (heatmap decoder), a Head Tracker (box decoder), and a Presence Predictor (in/out decoder).
The training objective is a joint multi-task objective
with Ltotal = Lgaze + Lpresence + Lhead,
where Lgaze is a pixel-wise binary cross-entropy (BCE) loss,
Lpresence is a Focal Loss supervised with a binary label,
and Lhead is a linear combination of the L1 loss and the generalized IoU loss.
The authors developed a dataset called Gaze-Co: We created Gaze-Co, the first large-scale dataset for PGE, containing 120K samples sourced from the training set of GazeFollow, VisualAttentionTarget (VAT) and ChildPlay.
The data engine involves three stages: (1) data alignment and filter; (2) concept generation; (3) verification.
Concept generation is done with a production Vision Language Model (VLM) accessed through API (Gemini 2.5 Pro)
and verification uses an Multi-modal Large Language Model (MLLM)-first, human-in-the-loop verification workflow.
In experiments, GazeAnywhere achieves state-of-the-art performance on multiple PGE benchmarks, setting a strong baseline for this new problem even on a difficult out-of-domain, real-world clinical dataset.
The main results table shows GazeAnywhere-DINOv3-L achieving the best performance on all metrics across GazeFollow-Concept, VAT-Concept, ChildPlay-Concept, and the OOD Child-SC dataset, compared to baselines pairing gaze models (ViTGaze, Sharingan, Gaze-LLE) with open-vocabulary detectors (GroundingDINO-B, LLMDet-L, OWLv2-L, RexSeek).
Ablation studies show: text-based prompting achieves performance on par with visual prompting
and the subject's appearance and pose description are the most critical components for the PGE task.
The loss ablation shows the presence loss only supports the auxiliary in/out prediction, not help gaze estimation. The head loss, however, improves both the gaze target estimation and the target presence prediction.
Encoder comparison shows the DINOv3-based model achieves the best performance on nearly all metrics.
The paper also demonstrates a real-world application: we developed the GazeAnywhere Agent... This system uses a central MLLM (Gemini 2.5) that leverages GazeAnywhere as a specialized tool to solve advanced user queries.
In testing on 10 real-world videos, the GazeAnywhere Agent demonstrated significantly better performance than a raw, single MLLM solution
for gaze shift and eye contact calculations.
The authors conclude: "We present GazeAnywhere, a system that enables interactive human gaze target estimation using flexible, open-vocabulary text prompts to identify the subject. Our principal contributions include introducing the novel Promptable Gaze Target Estimation (PGE) task and Gaze-Co benchmark, proposing a tailored transformer-based detector and learning objective, and developing a human-and-AI-in-the-loop data engine to adapt existing datasets. GazeAnywhere achieves state-of-the-art results in Gaze-Co benchmark, and its robustness is further validated on a challenging out-of-domain (OOD) dataset of child social communication videos."
Improvements for AI systems
Improvements to AI Systems:
-
Unified, Promptable Gaze Analysis Pipeline: Replace brittle multi-stage pipelines (head detection → pose estimation → gaze mapping) with a single end-to-end transformer that accepts natural language or point prompts (e.g.,
the woman in the blue coat
orperson at [0.3, 0.6]
) to directly output gaze heatmaps, head bounding boxes, and in/out-of-frame presence. This eliminates error cascading from upstream detectors. -
Open-Vocabulary Subject Grounding via Cross-Modal Fusion: Integrate a frozen vision encoder (e.g., DINOv3) and a frozen text encoder (e.g., CLIP-style) into a transformer detector that learns joint visual-textual representations. This enables the system to resolve arbitrary, unseen subject descriptions (appearance, pose, location) without retraining, unlike fixed-class detectors.
-
Auxiliary Task Co-Learning for Robustness: Add two auxiliary prediction heads—a Head Tracker (box regression with L1 + GIoU loss) and a Presence Predictor (binary focal loss)—to the main gaze heatmap head. The head loss improves gaze accuracy and presence prediction, while presence prediction provides explicit uncertainty for out-of-frame gaze cases, making the system more reliable in real-world scenes.
-
Scalable Synthetic Data Generation with Human-AI Verification: Use a production VLM (e.g., Gemini 2.5 Pro) to automatically generate diverse, concept-based text annotations from existing gaze datasets (GazeFollow, VAT, ChildPlay), then filter and verify via an MLLM-first, human-in-the-loop workflow. This creates 120K high-quality aligned pairs, enabling training without manual annotation.
-
Agentic Integration for Complex Queries: Embed the gaze model as a specialized tool inside a central MLLM agent (e.g., Gemini 2.5). The agent can decompose user queries like
Did the child look at the therapist during the last 30 seconds?
into calls to GazeAnywhere for gaze shifts and eye contact, then synthesize answers—outperforming a raw MLLM on real-world video analysis.
What the Improved AI System Can Do:
-
Interactive Gaze Analysis in Unconstrained Images: Given a natural language description or a click-point, the system instantly identifies the correct subject and estimates their gaze target (where they are looking) in a single forward pass, even in crowded scenes with multiple people.
-
Robust Performance on Out-of-Domain Data: Works on clinical child social-communication videos, where traditional detectors fail due to unusual poses, occlusion, or domain shift—thanks to promptable grounding and auxiliary head/presence tasks.
-
Zero-Shot Generalization to New Subjects: Handles arbitrary, unseen descriptions (e.g.,
the person holding a red cup
orthe child in the striped shirt
) without any fine-tuning, because text prompts are mapped to visual features via frozen encoders. -
Explainable and Uncertainty-Aware Outputs: Provides not only gaze heatmaps but also head location and a binary in/out-of-frame flag, allowing downstream systems to reason about gaze confidence and handle cases where the target is outside the image.
-
Automated Video Understanding via Agentic Reasoning: Answers complex temporal queries (e.g.,
How many times did the speaker look at the audience?
) by combining the gaze tool with an MLLM’s reasoning, enabling applications in social robotics, autism therapy monitoring, and human-robot interaction.
Abstract
Estimating human gaze targets from images in-the-wild is an important and formidable task. Existing approaches primarily employ brittle, multi-stage pipelines that require explicit inputs, like head bounding boxes and human pose, in order to identify the subject of gaze analysis. As a result, detection errors can cascade and lead to failure. Moreover, these prior works lack the flexibility of specifying the gaze analysis task via natural language prompting, an approach which has been shown to have significant benefits in convenience and scalability for other image analysis tasks. To overcome these limitations, we introduce the Promptable Gaze Target Estimation (PGE) task, a new end-to-end, concept-driven paradigm for gaze analysis. PGE conditions gaze prediction on flexible user text or visual prompts (e.g., "the boy in the red shirt" or "person in point [0.52, 0.48]") to identify a specific subject for gaze analysis. This approach integrates subject localization with gaze estimation, and eliminates the rigid dependency on intermediate analysis stages. We develop a scalable data engine to generate Gaze-Co (Gaze Estimation with Concepts), a dataset and benchmark of 120K high-quality, prompt-annotated image pairs. We also propose GazeAnywhere, the first model designed for PGE. GazeAnywhere uses a transformer-based detector to fuse features from frozen encoders and simultaneously solves subject localization, in/out-of-frame presence, and gaze target heatmap estimation. GazeAnywhere achieves state-of-the-art performance on multiple PGE benchmarks, setting a strong baseline for this new problem even on a difficult out-of-domain, real-world clinical dataset. GazeAnywhere is open-sourced in github.com/IrohXu/GazeAnywhere.
Sources
- Qwen2.5-VL Technical Report
- Perception Encoder: The best visual embeddings are not at the output of the network
- SAM 3: Segment Anything with Concepts
- Meta CLIP 2: A Worldwide Scaling Recipe
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- GazeDETR: Gaze Detection using Disentangled Head and Gaze Representations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- DINOv3
- SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
- GazeCLIP: Enhancing Gaze Estimation Through Text-Guided Multimodal Learning
- Demystifying CLIP Data
- ContextAgent: Context-Aware Proactive LLM Agents with Open-World Sensory Perceptions
- Advancing Complex Video Object Segmentation via Progressive Concept Construction
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models