Diffusion models for eye-gaze trajectory generation using position and velocity representations
cs.CV, cs.AI, cs.NE
Submitted: 2026-09-01
Updated: 2026-09-01
Comments: 14 pages, 7 figures
License: http://creativecommons.org/licenses/by/4.0/
The gist: Eye-tracking data are expensive to collect, requiring specialized hardware and controlled laboratory conditions, and difficult to share because of privacy constraints.
Terminology
Abstract
Eye-tracking data are expensive to collect, requiring specialized hardware and controlled laboratory conditions, and difficult to share because of privacy constraints. We address this using two complementary denoising diffusion probabilistic models (DDPMs) for unconditional generation of eye-gaze dynamics from visual-search data. Both use an identical FiLM-conditioned one-dimensional U-Net with self-attention (19.35,M parameters), trained on 8,s sliding-window sequences from 28 participants. One model generates raw two-dimensional gaze-position sequences, while the other generates two-component velocity sequences; each uses representation-specific preprocessing, training settings, data partitions, and evaluation protocols. Both are evaluated across three independent training seeds, with aggregated metrics reported as mean, plus or minus,SD. The position-space model achieves a mean Jensen-Shannon (JS) divergence of 0.016 plus or minus0.004 across nine kinematic features, with the highest feature-wise mean below 0.030, fixation duration within 2% of real data, and a Fr'echet Gaze Distance more than an order of magnitude below statistical and Markovian baselines. Under a Train-on-Synthetic-Test-on-Real protocol, synthetic-only training achieves R 2=0.66 plus or minus0.02, or 82.7% of the real-data R squared point estimate. The velocity-space model achieves a mean JS divergence of 0.0065 across velocity components, speed, log-speed, and turning angle, with a maximum of 0.015 plus or minus0.005. Reconstructed path length is less accurate (0.21 plus or minus0.02 versus 0.03 plus or minus0.01 in position space), although the protocols differ. Overall, unconditional diffusion captures local gaze kinematics and short-range temporal and directional structure, while long-range properties such as saccade counts and cumulative path geometry remain targets for future conditioned models.
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models