HAGI++: Head-Assisted Gaze Imputation and Generation

arXiv:2511.02468 · cs.HC, cs.CV · Submitted 2025-11-04 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "HAGI++: Head-Assisted Gaze Imputation and Generation".

Jane: HAGI++ introduces a novel multi-modal diffusion model that leverages the close coordination between eye and head movements to perform gaze data imputation and generation in real-world scenarios.

Tom: First, who's behind it and why it matters.

Title and authors: Tom: So, we're diving into the paper titled "HAGI++: Head-Assisted Gaze Imputation and Generation." It sounds like they've put a lot of focus on using the relationship between where your eyes are and how your head moves to fill in missing gaze data. Jane, what do you think is the big picture idea behind that title?

Jane: Well, Tom, it’s basically saying they aren't just looking at the eye data in isolation anymore; they're using the motion of the head as a guide to make things whole again. It simplifies the concept by showing how these two movements are tightly coupled in real life.

Lu: I think what’s really interesting here is that they’re treating this as a multi-modal diffusion model, which suggests a sophisticated way to handle these kinds of complex temporal dependencies within the data structure.

Meng: From an engineering standpoint, that means we're not just looking at one stream of data; we have to build a system that can manage and fuse those different signals coherently during the imputation process.

Lalam: I see this as a huge step for how we train generative models; if the model understands these physical coordination patterns, it should become much better at creating realistic outputs in general.

Tom: Exactly, Lalam, and that's what makes HAGI++ so compelling—it moves beyond just filling gaps to actually generating things that look natural. Does this mean we can finally get truly consistent gaze sequences?

Jane: It does; the paper suggests it tackles a major issue where simply discarding data creates unnatural jumps in movement. We’re looking at how this approach addresses that continuity problem.

Lu: They explicitly state that they exploit the fact that eye-head coordination is close, which is a strong physical assumption they are leveraging within their model architecture.

Meng: That reliance on physical assumptions is something we have to be careful about when we try to scale this up in a practical system.

Lalam: If it can produce more realistic trajectories, that’s fantastic for creating synthetic data that looks genuinely human and not just mathematically smoothed.

The paper's summary: Tom: So, let's look at what HAGI++ actually does based on the summary of "HAGI++: Head-Assisted Gaze Imputation and Generation." They propose a novel multi-modal diffusion model designed specifically for filling in missing gaze data by using time-aligned head movements as an auxiliary input.

Jane: That means instead of just trying to guess the missing points on their own, they condition the entire process on what the head is doing at that same moment, which should provide a much stronger context.

Lu: The core idea is exploiting those "close coordination between eye and head movements" to learn these cross-modal dependencies between the gaze representation and the head motion representation.

Meng: So, when we look at their methodology, it sounds like they're projecting both modalities into a shared latent space of dimension sixty-four using an MLP before feeding them into the main diffusion process. That’s a specific technical detail I want to dig into later.

Lalam: From my perspective as an AI, conditioning on multiple relevant inputs is always a win; it gives the model more constraints to follow during the generation steps, which leads to better control over the output space.

Tom: Right, and they structure this using a conditional score-based diffusion model for time-series imputation. The forward process adds noise to the missing parts, and then the reverse process iteratively denoises it using a learned deep learning model.

Jane: That iterative denoising is what allows it to progressively clean up the noisy data into something that looks like real gaze movement, rather than just a simple guess.

Lu: The architecture itself has several transformer blocks, each with self-attention for temporal dependencies and cross-attention layers specifically designed to model that coordination between gaze and body movements, like eye-head or even eye-hand-head.

Meng: That cross-attention mechanism sounds computationally intensive; we need to make sure the complexity of those attention operations scales well when we start adding more body movements later.

Lalam: The fact that it uses a transformer stack means it can capture very long-range temporal relationships, which is crucial for modeling continuous motion like gaze.

The paper's improvements: Tom: Now we’re talking about the specific improvements they claim in "HAGI++: Head-Assisted Gaze Imputation and Generation." They point out that traditional methods like classical interpolation or even some deep learning baselines struggle to accurately replicate the velocity profile of real human gaze data.

Jane: So, their main improvement is moving away from those older techniques because interpolation often fails to capture how fast things are actually moving in natural movement patterns.

Lu: They emphasize that their method achieves lower mean angular error compared to other approaches across various missing data ratios, which is a direct measure of accuracy in terms of the final gaze position prediction.

Meng: Lower angular error sounds good on paper, but I wonder if that translates to something useful when we're dealing with complex, noisy real-world data streams.

Lalam: It’s about achieving a better match for what real human eye movements actually look like in terms of direction and speed. That kind of fidelity is really important for any realistic simulation or analysis.

Tom: And they also look at the velocity distribution using Jensen–Shannon divergence, showing that their imputed gaze velocity distributions more closely resemble those from real human gaze movements, which is a different way to measure success.

Jane: So, not only are the positions accurate but the way the movement flows is also much closer to what we see in actual human eyes moving. That’s a big win for plausibility.

Lu: Furthermore, their ablation studies showed that head rotation consistently outperforms head translation in terms of imputation accuracy, suggesting that rotational components are more dominant drivers of gaze prediction fidelity than simple linear movement.

Meng: That finding about head rotation being more important than translation gives us some practical insights into which sensor data we should prioritize when building a system.

Lalam: And if we add wrist information, the results improved significantly under higher missing ratios, showing that the eye-head-hand coordination becomes more useful over longer durations.

Conclusion: Tom: Okay, so to wrap up on "HAGI++: Head-Assisted Gaze Imputation and Generation," it seems they've developed a robust framework that successfully uses multi-modal input to create more realistic gaze data compared to previous state-of-the-art methods.

Jane: We’re looking at a solution that handles missing data by leveraging the coordination between eye and head movements, which is what makes this approach so effective for making those gaps look plausible.

Lu: The implications for XR environments are significant because it opens up the possibility of generating synthetic gaze data directly from head tracking information when eye-tracking hardware isn't present.

Meng: For implementation, the paper flags that they still need to ensure the complexity of their transformer stack doesn't lead to too much overhead in deployment compared to simpler interpolation methods.

Lalam: I think the biggest impact is on culture because it pushes us toward building AI systems that are inherently more context-aware and physically grounded, which can make our applications feel much more intuitive for users.

CHUHAN JIAO, ZHIMING HU, ANDREAS BULLING

University of Stuttgart

cs.HC, cs.CV

Submitted: 2025-11-04

Updated: 2026-09-28

Importance score: 92/100

The gist: HAGI++ introduces a novel multi-modal diffusion model that leverages the close coordination between eye and head movements to perform gaze data imputation and generation in real-world scenarios.

Key concepts

Multi-modal Diffusion Model
A type of deep learning model that learns to generate or fill in missing data by adding noise and then iteratively removing it. HAGI++ uses this approach to handle complex gaze data by combining visual gaze information with motion data from the head and body.
FiLM Layer
A mechanism used within the neural network that allows for feature-wise linear modulation. In HAGI++, this layer is used to effectively combine different types of motion features (like head and gaze) at various levels of the model, ensuring continuous fusion of this auxiliary information.
Cross-modal Dependency
The ability of a model to understand how two different types of data relate to each other. HAGI++ specifically learns the dependency between eye movements and head movements, allowing it to use head motion as a guide to accurately predict the missing gaze path.

Terminology

Summary

HAGI++ introduces a novel multi-modal diffusion model that leverages the close coordination between eye and head movements to perform gaze data imputation and generation in real-world scenarios. This method is significant because it exploits inherent correlations between eye and head movements, leading to more realistic gaze trajectories and superior performance across various data loss conditions compared to conventional interpolation or deep learning baselines.

How it works

HAGI++ is a transformer-based diffusion model designed for gaze imputation that explicitly incorporates time-aligned head movements as an auxiliary modality. The core idea is to exploit the close coordination between eye and head movements by conditioning the diffusion process on observed gaze data and time-aligned head motion. This approach allows the model to learn cross-modal dependencies between eye and head representations.

The framework is designed to be extensible, allowing it to incorporate additional body movements. Specifically, HAGI++ employs a hybrid feature-fusion mechanism based on FiLM (feature-wise linear modulation) to effectively combine head, gaze, and other body motion features across multiple levels. The input modalities are projected into a shared latent space of dimension 64 using a Multilayer Perceptron (MLP) followed by a GELU activation. This results in unified gaze tensor G and motion context tensor B, which serve as auxiliary inputs to guide the imputation process.

Key Methodological Components

The method is structured around a conditional score-based diffusion model for time-series imputation (CSDI), adapted for gaze data. The forward process progressively adds noise to the target part of the gaze sequence, while the reverse process iteratively denoises it using a learned denoising deep learning model. The training objective minimizes the difference between the prediction and the actual added noise at each diffusion time step, as defined by Equation 9.

The architecture consists of a stack of N transformer blocks. Each block includes:

  1. A self-attention mechanism to capture temporal dependencies and spatial relationships between observed and missing gaze samples (Equation 10).

  2. A cross-attention layer to model the coordination between gaze and body movements (e.g., eye–head or eye–hand–head coordination) by attending to the motion context tensor B (Equation 11).

  3. A FiLM layer that utilizes a skip fusion mechanism to re-inject motion information at each Transformer block, ensuring continuous fusion of auxiliary motion context with the gaze representation (Equation 12).

Input and Data Preparation

The input data is structured as follows:

Gaze Data:

The gaze sequence X is represented as a sequence of (pitch, yaw) vectors. It is divided into an observed segment xc0 and a target segment xta0 using a randomly generated binary observation mask M.

Head Movements:

Head movements H are captured by mobile eye trackers and are processed into 12-dimensional vectors through Fourier positional encoding to obtain the encoded head movements H ∈ R L×12.

Auxiliary Body Movements (Optional):

When available, wrist or hand movements W are incorporated. These relative transformations between the eye tracker and the wearable device are also projected into a shared latent space, forming part of the motion context tensor B.

Evaluation and Results

HAGI++ was evaluated on three large-scale egocentric datasets: Nymeria, Ego-Exo4D, and HOT3D. The performance was assessed using two primary metrics:

  1. Mean Angular Error (MAE): Measures the angular difference between predicted and ground truth gaze vectors (Equation 13). HAGI++ consistently achieved lower MAE across all missing data ratios compared to baselines, achieving an average improvement of 24.45% compared with the previous state-of-the-art method.

  2. Jensen–Shannon divergence (JS): Measures the similarity between the velocity distribution of imputed and real human gaze movements (Equation 14). HAGI++ produces gaze velocity distributions that more closely resemble real human gaze movements, achieving performance on par with prior state-of-the-art methods.

Ablation and Coordination Insights

Ablation studies revealed the independent contributions of different motion components. The results showed that head rotation consistently outperforms HAGI++ translation in terms of imputation accuracy, suggesting that head rotation plays a more dominant role in gaze imputation performance than head translation. Furthermore, incorporating wrist information was found to improve performance significantly under higher missing ratios (90% and 100% data loss), demonstrating the synergistic relationship between eye, head, and hand movements. This suggests that over longer durations, eye–hand coordination begins to contribute more substantially to gaze prediction.

Conclusion

HAGI++ is a robust and practical solution for enhancing gaze data quality in real-world and XR environments by effectively modeling multi-modal eye–head–hand coordination.

Improvements for AI systems

As a fastidious researcher, I have analyzed the provided paper, HAGI++: Head-Assisted Gaze Imputation and Generation, which introduces a novel multi-modal diffusion model for gaze data.

Here are the specific improvements that can be made to existing AI systems, based on the capabilities demonstrated by HAGI++, and what those improved systems can achieve:


  1. The integration of cross-modal dependencies between eye and head representations using a transformer-based diffusion model.

  2. The ability to condition the diffusion process not only on gaze but also on time-aligned head movements, and optionally, auxiliary body movements (e.g., wrist motions).

  3. The use of a hybrid feature fusion mechanism based on FiLM (Feature-wise Linear Modulation) blocks to effectively combine head, gaze, and other body motion features across multiple levels throughout the denoising process.

Specific improvements to AI systems:

  1. A robust system capable of accurately imputing missing gaze data in real-world mobile eye-tracking scenarios by explicitly exploiting the inherent biomechanical correlation between head and eye movements.

  2. An XR/VR gaze generation system that can synthesize highly realistic, biologically plausible gaze trajectories (including velocity distributions) from only head and commodity wearable device motion signals, overcoming the limitation of requiring full-body motion capture.

  3. A data processing pipeline that can dynamically incorporate auxiliary inputs (like wrist movements from smartwatches) to enhance imputation accuracy in high data-loss scenarios, demonstrating a synergistic relationship between eye-head and eye-hand coordination for superior gaze reconstruction.

What the improved AI system can do:

  1. It will be able to provide continuous, temporally complete gaze trajectories even when the input tracking signal is interrupted (e.g., due to blinks or occlusion), maintaining a high degree of biological plausibility in human eye movement dynamics.

  2. It can generate synthetic gaze data for applications where real eye-tracking hardware is unavailable (e.g., XR headsets without built-in eye-tracking) by leveraging head tracking data, enabling the creation of realistic visual feedback for virtual environments.

  3. It can enable more accurate gaze-based analysis and interaction in complex environments by modeling the coordination between head orientation and gaze shifts, leading to better predictions of user intent or cognitive states during activities like cooking or object manipulation.

  4. It can serve as a powerful post-processing tool for machine learning models that require complete time-series data, effectively filling in gaps with human-like data rather than discarding the entire sequence.

Sources

Related papers