Real-time Appearance-based Gaze Estimation for Open Domains

arXiv:2603.26945 · cs.CV · Submitted 2026-03-27 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "Real-time Appearance-based Gaze Estimation for Open Domains".

Tom: Appearance-based gaze estimation (AGE) models often fail in practical, unconstrained scenarios due to limited image diversity and inconsistent label fidelity across different datasets, particularly along the pitch axis.

Jane: First, who's behind it and why it matters.

Paper summary: Lu: To conclude, the paper "Real-time Appearance-based Gaze Estimation for Open Domains" focuses on overcoming the generalization gap in appearance-based gaze estimation by addressing limited image diversity and inconsistent label fidelity across datasets.

Jane: They propose a framework that expands this image manifold through augmentation and reformulates the task as a multi-task learning problem incorporating auxiliary objectives like discretized classification, segmentation, and contrastive learning.

Tom: The authors demonstrate that this approach enhances generalization without needing additional human-annotated data by creating new benchmarks like RealGaze and ZeroGaze to rigorously test its robustness against real-world conditions.

Meng: The implications for practical AI systems are significant because they show a pathway to deploy these gaze estimation tools effectively in environments where they aren't perfectly controlled, like those involving wearables or poor lighting.

Lalam: From a cultural perspective, if we can make gaze estimation more robust across diverse visual conditions, it could mean better interaction design for assistive technologies and personalized AI interfaces everywhere.

Lu: It also shows that even when dealing with anisotropic label deviation in datasets like DN and DC, selective loss application can successfully shift supervision to relative manifold alignment.

Jane: Overall, the work suggests that by making appearance-based gaze estimation more aware of its input variations and internal inconsistencies, we can build more reliable AI systems for this specific vision task.

Tom: It’s a solid piece of research showing how to make appearance-based gaze estimation more resilient in messy, unconstrained environments without needing constant new labeling effort.

Conclusion: Tom: So, we're wrapping up our discussion on this paper about real-time appearance-based gaze estimation for open domains, and I want to quickly recap what we've covered before we get into the big picture stuff.

Jane: Exactly, Tom. We looked at how they tackle those common problems in gaze estimation by adding extra layers of learning and new data sources to make the model much more reliable in the real world.

Lu: I found the way they reframe it as a multi-task learning problem really interesting because it opens up so many creative avenues for how we might approach vision tasks generally.

Meng: From an engineering standpoint, I'm still thinking about how much training data these new augmentation techniques actually need to handle in practice, and that's a real hurdle.

Lalam: I think what really stands out is the way the model learns to ignore irrelevant factors like glasses or masks through those contrastive terms; that level of feature robustness is impressive for an AI system.

Tom: Right, that robustness is key when you think about how we can actually deploy these systems in unpredictable settings. So, let's talk about what this whole thing means for the future and the world outside of the lab.

Jane: It really does. This research suggests that we can move gaze estimation out of perfectly controlled environments and into everyday applications where things are messy and unpredictable.

Lu: Imagine assistive technology that works reliably whether someone is wearing a specific type of eyewear or if there's some weird lighting in the room; that’s a massive potential shift for accessibility.

Meng: I wonder how this translates to the startup world—can we actually implement these lightweight models on edge devices like phones without sacrificing too much accuracy? That practical deployment aspect is something I'm thinking about.

Lalam: If this technology becomes more robust, it could fundamentally change how AI interacts with us, making interfaces feel much more intuitive and less frustrating to use for everyone.

Tom: It sounds like we're looking at a future where gaze interaction feels much more natural and less prone to those annoying errors we see in current systems. So, what are the authors saying about the next steps for this line of research?

Huawei Technologies Canada · University of Toronto

cs.CV

Submitted: 2026-03-27

Updated: 2026-10-01

Code: https://github.com/black-forest-labs/flux

Importance score: 90/100

The gist: Appearance-based gaze estimation (AGE) models often fail in practical, unconstrained scenarios due to limited image diversity and inconsistent label fidelity across different datasets, particularly

Key concepts

Automated Augmentation Pipeline
This pipeline automatically creates more training data by applying various techniques such as synthesizing eyeglasses, masks, and changing lighting. It also uses noise models to simulate sensor errors and portrait matting for background replacement with random indoor scenes, effectively expanding the image variety the model sees.
Multi-task Learning (MTL)
Instead of just predicting gaze coordinates, the model is trained on several related tasks simultaneously. This includes classification of discretized gaze ranges and segmentation of eyes. These auxiliary objectives help regularize the feature space and prevent certain types of errors, leading to more robust predictions.
Multi-view Supervised Contrastive (SupCon) Learning
This technique uses four specific contrastive terms to pull features from positive pairs closer together while pushing negative pairs apart. It enforces invariance by ensuring the model learns features that are consistent regardless of factors like head pose, whether glasses are present, or if a mask is worn.

Terminology

Summary

Appearance-based gaze estimation (AGE) models often fail in practical, unconstrained scenarios due to limited image diversity and inconsistent label fidelity across different datasets, particularly along the pitch axis. This paper proposes a robust AGE framework that enhances generalization without requiring additional human-annotated data by expanding the image manifold through augmentation and reformulating gaze regression as a multi-task learning problem incorporating auxiliary objectives.

How it works

The proposed framework addresses generalization gaps through three core strategies: expanding image diversity, reformulating the task, and curating new benchmarks. First, an automated augmentation pipeline is introduced to expand the image manifold via an ensemble of augmentation techniques, including synthesis of eyeglasses, masks, and varied lighting conditions. This includes portrait matting for background replacement with random indoor scenes and heuristic noise models to approximate sensor-level artifacts.

Second, the framework reformulates AGE as a multi-task learning problem that supplements regression with several auxiliary objectives to regularize the feature space. These objectives include:

  1. Discretized label classification, which involves partitioning the gaze range into an nϕ × nψ grid and supervising both a cross-entropy loss and an L1 regression loss to prevent mean-collapse.

  2. Eye and iris segmentation, using MediaPipe landmarks to generate ground truths via a pipeline that includes brightness thresholding and morphological operations, with supervision via the Dice loss (Lseg) to anchor the representation of ocular appearance.

  3. Multi-view Supervised Contrastive (SupCon) learning, which introduces four specific terms: a pitch contrastive term (L Sϕ), a dataset-invariance term (L S D), a glasses-invariance term (L S g), and a mask-invariance term (L S m). These terms pulling features from positive pairs closer together than negative pairs to enforce invariance to gaze-irrelevant factors.

Data Constraints and Anisotropy

The paper identifies two systemic constraints in existing datasets: the Diversity-Fidelity Tradeoff, where standard labeling restricts identity diversity and environmental variation, and Anisotropic Inter-dataset Label Deviation. A perceptual study reveals that pitch labels exhibit substantially lower perceptual consistency than yaw across all dataset combinations, suggesting that pitch labels are more susceptible to systematic errors during the PoG-to-3D mapping process. Consequently, the framework selectively discards Lreg and Lclf for pitch labels from datasets like DN and DC where fidelity is lower, instead using a pitch-aware SupCon loss L Sϕ to shift supervision from absolute coordinates to relative manifold alignment.

Model Architecture and Training Objectives

The model architecture is based on a lightweight MobileNet-v2 enhanced with Coordinate Attention (CA) for an expanded receptive field. The overall training objective is defined as:

L = Lreg + λclfLclf + λsegLseg + λDL S D + λϕL Sϕ + λgL S g + λmL S m.

The framework utilizes a streamlined architecture during inference with a MobileNet-based lightweight model that achieves generalization performance competitive with the state-of-the-art UniGaze-H while utilizing less than 1% of its parameters.

Validation and Evaluation

To rigorously validate the approach, two new benchmark datasets are curated: RealGaze, which emulates real-world application using a 13-inch tablet with varied wearables and lighting conditions, and ZeroGaze, a synthesized dataset generated via Flux.1 to isolate the impact of head pose and gaze variance by setting ground truth labels to zero. Evaluation on these benchmarks demonstrates that the proposed method maintains superior robustness and manifold consistency compared to significantly larger baselines, maintaining concentrated distributions around the zeropoint on ZeroGaze, effectively seeing through the occlusions that confuse standard regressors. Furthermore, training-free personalized calibration is proposed as a lightweight mechanism to bridge the gap between laboratory performance and real-world reliability.

Ablation and Performance Gains

A series of ablations confirm the contribution of each component:

- Augmentation and SupCon data augmentation is central to generalization performance.

- The segmentation task further improves accuracy in occlusion-heavy sessions by anchoring the model’s attention to the ocular region.

The results show that our augmentation pipeline and multi-task supervision substantially improve generalization to challenging conditions. In cross-dataset experiments, the MobileNet model outperforms heavy baselines such as UniGaze-H while using under 1% of its parameters. The training-free calibration method shows that a single calibration point provides the most significant performance leap, particularly in mitigating the pitch bias identified in our ZeroGaze experiments. In RealGaze evaluation, the method "outperforms the competition by a large margin, particularly in dY and∥d∥2.

Improvements for AI systems

As a fastidious and diligent researcher, I have analyzed the proposed framework in Real-time Appearance-based Gaze Estimation for Open Domains. The core contribution is a robust, lightweight Age Estimation (AGE) model that achieves superior generalization in unconstrained, real-world scenarios by addressing data diversity and label inconsistency through multi-task learning and advanced augmentation.

Here are the specific improvements to AI systems achievable by implementing this framework:


The proposed system can be implemented as a highly efficient, end-to-end gaze tracking pipeline suitable for mobile devices. The following specific capabilities are enabled:

  1. [Image/Video] Gaze Estimation in Unconstrained Environments:

  2. [Robustness to Occlusion and Wearables]:

  3. [Real-time Low-Latency Tracking]:

  4. [Cross-Domain Generalization]:

  5. [Personalized, Training-Free Calibration for High Precision]:

Detailed technical specifications of what the improved AI system can do:

  1. A lightweight MobileNet-v2 based model can accurately estimate 3D gaze direction (pitch and yaw) from RGB images in challenging real-world conditions, including those involving eyeglasses and facial masks, achieving performance competitive with heavy state-of-the-art models like UniGaze-H while utilizing under 1% of the parameters.

  2. The system will maintain high accuracy even under poor or directional lighting conditions (side-lighting), as it is explicitly trained to focus on geometric ocular structures rather than illumination artifacts, a capability demonstrated by its superior performance in Side-Lit sessions on the RealGaze benchmark.

  3. The model can perform real-time gaze tracking directly on mobile devices due to its lightweight architecture, enabling high-fidelity tracking with low latency (as opposed to heavy backbones) suitable for practical applications like augmented reality (AR) overlays or interactive user interfaces.

  4. The system exhibits strong cross-domain generalization, meaning it performs reliably on unseen datasets and in novel scenarios by utilizing an automated augmentation pipeline that synthesizes eyeglasses, masks, and varied lighting conditions to expand the image manifold.

  5. The system can achieve high precision (sub-millimeter error) in gaze estimation through a training-free personalized calibration mechanism. This allows the AI to adapt its inherent pitch bias (caused by inter-dataset label deviation) to a specific user's physiological variance in real-time or via a brief, low-overhead calibration process, significantly mitigating systematic errors that plague uncalibrated models.

Sources

Related papers