Gaze Target Estimation Anywhere with Concepts

arXiv:2608.11367 · cs.CV, cs.AI · Submitted 2026-08-11 · Read on arXiv

Xu Cao, Houze Yang, Vipin Gunda, Zhongyi Zhou, Tianyu Xu, Adarsh Kowdle, Inki Kim, James M. Rehg

University of Illinois Urbana-Champaign · Google

cs.CV, cs.AI

Submitted: 2026-08-11

Updated: 2026-08-13

Comments: CVPR 2026 Code and Benchmark are aviliable at https://github.com/IrohXu/GazeAnywhere and https://huggingface.co/datasets/IrohXu/Gaze-Co-Benchmark

Code: https://github.com/IrohXu/GazeAnywhere

License: http://creativecommons.org/licenses/by-nc-sa/4.0/

Importance score: 75/100

The gist: This paper introduces the Promptable Gaze Target Estimation (PGE) task and the GazeAnywhere model, a new end-to-end, concept-driven paradigm for gaze analysis.

Terminology

Summary

This paper introduces the Promptable Gaze Target Estimation (PGE) task and the GazeAnywhere model, a new end-to-end, concept-driven paradigm for gaze analysis. The authors state: "We introduce the Promptable Gaze Target Estimation (PGE) task, a new end-to-end, concept-driven paradigm for gaze analysis. PGE conditions gaze prediction on flexible user text or visual prompts (e.g., 'the boy in the red shirt' or 'person in point [0.52, 0.48]') to identify a specific subject for gaze analysis."

The paper's main contributions are summarized as: "(1) We define the Promptable Gaze Target Estimation (PGE) task, which extends the traditional gaze target estimation problem to an unconstrained, text promptable end-to-end paradigm. (2) We design a scalable data engine to generate 120K high quality PGE training annotations consisting of subject text to gaze alignment data pairs. (3) We introduce GazeAnywhere, the first promptable concept-driven gaze target estimation model."

The authors motivate the work by noting that "Existing approaches primarily employ brittle, multi-stage pipelines that require explicit inputs, like head bounding boxes and human pose, in order to identify the subject of gaze analysis. As a result, detection errors can cascade and lead to failure. Moreover, these prior works lack the flexibility of specifying the gaze analysis task via natural language prompting."

The GazeAnywhere architecture is described as follows: "The model consists of a frozen image encoder, a frozen text encoder to proceed visual modality and text modality. A transformer-based detector is used to learn joint representations and map text prompt into the main gaze target estimation and auxiliary tasks. Specifically, it uses a frozen ViT, denoted ϕV (·), as our image encoder and a frozen text encoder, ϕT (·), which consists of a series of transformer blocks and a final linear layer. The model includes a head token for explicitly predict the head localization of the prompted subject and a Target Presence Token for the in/out-of-frame gaze target boolean prediction objective. The detector transformer fuse[s] these representations and refine[s] them for the gaze target estimation task and outputs a refined sequence of tokens which are passed to three distinct prediction heads": a Gaze Tracker (heatmap decoder), a Head Tracker (box decoder), and a Presence Predictor (in/out decoder).

The training objective is a joint multi-task objective with Ltotal = Lgaze + Lpresence + Lhead, where Lgaze is a pixel-wise binary cross-entropy (BCE) loss, Lpresence is a Focal Loss supervised with a binary label, and Lhead is a linear combination of the L1 loss and the generalized IoU loss.

The authors developed a dataset called Gaze-Co: We created Gaze-Co, the first large-scale dataset for PGE, containing 120K samples sourced from the training set of GazeFollow, VisualAttentionTarget (VAT) and ChildPlay. The data engine involves three stages: (1) data alignment and filter; (2) concept generation; (3) verification. Concept generation is done with a production Vision Language Model (VLM) accessed through API (Gemini 2.5 Pro) and verification uses an Multi-modal Large Language Model (MLLM)-first, human-in-the-loop verification workflow.

In experiments, GazeAnywhere achieves state-of-the-art performance on multiple PGE benchmarks, setting a strong baseline for this new problem even on a difficult out-of-domain, real-world clinical dataset. The main results table shows GazeAnywhere-DINOv3-L achieving the best performance on all metrics across GazeFollow-Concept, VAT-Concept, ChildPlay-Concept, and the OOD Child-SC dataset, compared to baselines pairing gaze models (ViTGaze, Sharingan, Gaze-LLE) with open-vocabulary detectors (GroundingDINO-B, LLMDet-L, OWLv2-L, RexSeek).

Ablation studies show: text-based prompting achieves performance on par with visual prompting and the subject's appearance and pose description are the most critical components for the PGE task. The loss ablation shows the presence loss only supports the auxiliary in/out prediction, not help gaze estimation. The head loss, however, improves both the gaze target estimation and the target presence prediction. Encoder comparison shows the DINOv3-based model achieves the best performance on nearly all metrics.

The paper also demonstrates a real-world application: we developed the GazeAnywhere Agent... This system uses a central MLLM (Gemini 2.5) that leverages GazeAnywhere as a specialized tool to solve advanced user queries. In testing on 10 real-world videos, the GazeAnywhere Agent demonstrated significantly better performance than a raw, single MLLM solution for gaze shift and eye contact calculations.

The authors conclude: "We present GazeAnywhere, a system that enables interactive human gaze target estimation using flexible, open-vocabulary text prompts to identify the subject. Our principal contributions include introducing the novel Promptable Gaze Target Estimation (PGE) task and Gaze-Co benchmark, proposing a tailored transformer-based detector and learning objective, and developing a human-and-AI-in-the-loop data engine to adapt existing datasets. GazeAnywhere achieves state-of-the-art results in Gaze-Co benchmark, and its robustness is further validated on a challenging out-of-domain (OOD) dataset of child social communication videos."

Improvements for AI systems

Improvements to AI Systems:

  1. Unified, Promptable Gaze Analysis Pipeline: Replace brittle multi-stage pipelines (head detection → pose estimation → gaze mapping) with a single end-to-end transformer that accepts natural language or point prompts (e.g., the woman in the blue coat or person at [0.3, 0.6]) to directly output gaze heatmaps, head bounding boxes, and in/out-of-frame presence. This eliminates error cascading from upstream detectors.

  2. Open-Vocabulary Subject Grounding via Cross-Modal Fusion: Integrate a frozen vision encoder (e.g., DINOv3) and a frozen text encoder (e.g., CLIP-style) into a transformer detector that learns joint visual-textual representations. This enables the system to resolve arbitrary, unseen subject descriptions (appearance, pose, location) without retraining, unlike fixed-class detectors.

  3. Auxiliary Task Co-Learning for Robustness: Add two auxiliary prediction heads—a Head Tracker (box regression with L1 + GIoU loss) and a Presence Predictor (binary focal loss)—to the main gaze heatmap head. The head loss improves gaze accuracy and presence prediction, while presence prediction provides explicit uncertainty for out-of-frame gaze cases, making the system more reliable in real-world scenes.

  4. Scalable Synthetic Data Generation with Human-AI Verification: Use a production VLM (e.g., Gemini 2.5 Pro) to automatically generate diverse, concept-based text annotations from existing gaze datasets (GazeFollow, VAT, ChildPlay), then filter and verify via an MLLM-first, human-in-the-loop workflow. This creates 120K high-quality aligned pairs, enabling training without manual annotation.

  5. Agentic Integration for Complex Queries: Embed the gaze model as a specialized tool inside a central MLLM agent (e.g., Gemini 2.5). The agent can decompose user queries like Did the child look at the therapist during the last 30 seconds? into calls to GazeAnywhere for gaze shifts and eye contact, then synthesize answers—outperforming a raw MLLM on real-world video analysis.

What the Improved AI System Can Do:

  • Interactive Gaze Analysis in Unconstrained Images: Given a natural language description or a click-point, the system instantly identifies the correct subject and estimates their gaze target (where they are looking) in a single forward pass, even in crowded scenes with multiple people.

  • Robust Performance on Out-of-Domain Data: Works on clinical child social-communication videos, where traditional detectors fail due to unusual poses, occlusion, or domain shift—thanks to promptable grounding and auxiliary head/presence tasks.

  • Zero-Shot Generalization to New Subjects: Handles arbitrary, unseen descriptions (e.g., the person holding a red cup or the child in the striped shirt) without any fine-tuning, because text prompts are mapped to visual features via frozen encoders.

  • Explainable and Uncertainty-Aware Outputs: Provides not only gaze heatmaps but also head location and a binary in/out-of-frame flag, allowing downstream systems to reason about gaze confidence and handle cases where the target is outside the image.

  • Automated Video Understanding via Agentic Reasoning: Answers complex temporal queries (e.g., How many times did the speaker look at the audience?) by combining the gaze tool with an MLLM’s reasoning, enabling applications in social robotics, autism therapy monitoring, and human-robot interaction.

Abstract

Estimating human gaze targets from images in-the-wild is an important and formidable task. Existing approaches primarily employ brittle, multi-stage pipelines that require explicit inputs, like head bounding boxes and human pose, in order to identify the subject of gaze analysis. As a result, detection errors can cascade and lead to failure. Moreover, these prior works lack the flexibility of specifying the gaze analysis task via natural language prompting, an approach which has been shown to have significant benefits in convenience and scalability for other image analysis tasks. To overcome these limitations, we introduce the Promptable Gaze Target Estimation (PGE) task, a new end-to-end, concept-driven paradigm for gaze analysis. PGE conditions gaze prediction on flexible user text or visual prompts (e.g., "the boy in the red shirt" or "person in point [0.52, 0.48]") to identify a specific subject for gaze analysis. This approach integrates subject localization with gaze estimation, and eliminates the rigid dependency on intermediate analysis stages. We develop a scalable data engine to generate Gaze-Co (Gaze Estimation with Concepts), a dataset and benchmark of 120K high-quality, prompt-annotated image pairs. We also propose GazeAnywhere, the first model designed for PGE. GazeAnywhere uses a transformer-based detector to fuse features from frozen encoders and simultaneously solves subject localization, in/out-of-frame presence, and gaze target heatmap estimation. GazeAnywhere achieves state-of-the-art performance on multiple PGE benchmarks, setting a strong baseline for this new problem even on a difficult out-of-domain, real-world clinical dataset. GazeAnywhere is open-sourced in github.com/IrohXu/GazeAnywhere.

Sources

Related papers