Gaze Estimation for Human-Robot Interaction: Analysis Using the NICO Platform
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "Gaze Estimation for Human-Robot Interaction".
Tom: The gist: This paper evaluates four state-of-the-art gaze estimation models in a shared workspace scenario using an annotated dataset collected with the NICO robotic platform,
Jane: First, who's behind it and why it matters.
Paper summary: Tom: So we're wrapping up our look at "Gaze Estimation for Human-Robot Interaction: Analysis Using the NICO Platform," and what they’re showing us is how current gaze estimation methods actually perform when a human and a robot are working together in the same space.
Jane: That’s right, Tom. The authors took some existing models and put them through a rigorous test using their own NICO robotic platform to see how accurate they are when a human and that robot are sharing the same table space <ref:2509.24001#pg1>.
Lu: This paper is looking at four state-of-the-art gaze estimation methods, testing them in a shared workspace scenario, which is pretty practical research <ref:2509.24001#pg3>.
Meng: The authors collected their own annotated dataset using the NICO platform for this evaluation, so it’s grounded in real robot interaction data rather than just old benchmarks <ref:2509.24001#pg3>.
Tom: They found that while the angular errors—how wrong the direction is—are pretty comparable to other general gaze estimation benchmarks, the distance errors are more limited in practice <ref:2509.24001#pg2>.
Jane: This means they focused on how accurate these models are when we need to know exactly where someone's eyes are looking on a table surface <ref:2509.24001#pg3>.
Lalam: For the cultural impact of this work, it tells us that for AI to really be useful in shared environments, it can't just guess where someone is looking; it needs a way to combine that limited spatial data with other context <ref:2509.24001#pg3>.
Tom: Right. So the paper concludes by summarizing that these methods are usable but have real limitations when we need high precision for a shared task scenario <ref:2509.24001#pg3>.
Jane: And they suggest that instead of relying solely on gaze for exact placement, we should consider using it for broader cues, like figuring out if the person is looking at the robot or just at the workspace around them <ref:2509.24001#pg3>.
Lu: It opens up a path where these estimation networks can be supplemented by other sensory inputs to fill in those gaps in understanding <ref:2509.24001#pg3>.
Meng: From an engineering standpoint, that makes sense; we don't need perfect gaze localization for everything; we need contextual awareness, and this paper points toward that approach <ref:2509.24001#pg3>.
Tom: And it leads us into what the authors suggest next, looking at how we can actually make these systems more robust by combining different types of data.
Conclusion: Tom: We’re wrapping up our look at "Gaze Estimation for Human-Robot Interaction: Analysis Using the NICO Platform," and what they really showed us is that these gaze estimation methods have a real spatial limit when we’re working with robots in shared spaces.
Jane: That’s right, Tom. The paper comes from researchers who used the NICO platform to test four different state-of-the-art gaze models in a table surface scenario <ref:2509.24001#pg1>.
Lu: So the focus here is on how these models handle real interaction, not just abstract data sets, which is pretty important for practical AI development <ref:2509.24001#pg3>.
Meng: The authors built their own dataset using NICO cameras, so they're grounding this evaluation in actual robot interaction data instead of just using old benchmarks <ref:2509.24001#pg3>.
Tom: They found that while the angle—the direction—they estimate is pretty good compared to other general tests, the actual location in space is where things get tricky <ref:2509.24001#pg3>.
Jane: That spatial uncertainty means they focused on how accurate these models are when we need to know exactly where someone’s eyes are landing on a table surface <ref:2509.24001#pg3>.
Lalam: For the cultural impact of this work, it tells us that for AI to really understand social cues in a shared environment, it can't just guess where someone is looking; it needs a way to combine that limited spatial data with other context <ref:2509.24001#pg3>.
Tom: Right. So the paper concludes by summarizing that these methods are usable but have real limitations when we need high precision for a shared task scenario <ref:2509.24001#pg3>.
Jane: And they suggest that instead of relying solely on gaze for exact placement, we should consider using it for broader cues, like figuring out if the person is looking at the robot or just at the workspace around them <ref:2509.24001#pg3>.
Lu: It opens up a path where these estimation networks can be supplemented by other sensory inputs to fill in those gaps in understanding <ref:2509.24001#pg3>.
Meng: For practical engineering, that tells us we don't need the gaze system to be perfect for everything; we just need it to give us a reliable starting point before we integrate other sensors for the final positioning <ref:2509.24001#pg3>.
Tom: So the paper boils down to this: gaze estimation is a useful tool, but in real-world robotics, you have to be realistic about the spatial resolution you're actually getting from it <ref:2509.24001#pg3>.
Jane: And that realism is key if we want to design systems that actually work well when humans and robots are collaborating in the same room <ref:2509.24001#pg3>.
Tom: Now, the next thing we’re going to look at is how these specific results change what engineers are building right now.
Faculty of Mathematics, Physics and Informatics, Comenius University · Italian Institute of Technology
cs.CV, cs.RO
Submitted: 2025-09-28
Updated: 2026-10-08
Code: https://github.com/kocurvik/nico_gaze
Importance score: 56/100
The gist: The gist: This paper evaluates four state-of-the-art gaze estimation models in a shared workspace scenario using an annotated dataset collected with the NICO robotic platform, finding that while
Key concepts
- Gaze Estimation
- This is the process of using cameras and computer networks to predict where a human is looking. In this study, it was used to understand social cues in human-robot interaction, helping systems interpret intentions and coordinate tasks by analyzing where people direct their attention.
- Shared Workspace Scenario
- The experiment took place when a human and a robot worked together on the same surface, like a table. This setup mimics real-world HRI situations where both entities are in close proximity, making gaze estimation relevant for understanding joint attention and task coordination.
- Angular Error
- This metric measures how far off the estimated direction of a person's gaze is from their actual gaze direction. The paper found that the models performed reasonably well in terms of this angle, suggesting they can capture general directional information accurately.
- Distance Error (in Plane)
- This metric calculates how far off the estimated point on a flat work surface is from the true target location. The study highlighted a significant limitation here, showing that while angles might be good, pinpointing exact locations on the table plane remains less accurate.
Terminology
Summary
The gist: This paper evaluates four state-of-the-art gaze estimation models in a shared workspace scenario using an annotated dataset collected with the NICO robotic platform, finding that while angular errors are comparable to general benchmarks, distance errors are limited to a median of 16.48 cm for the best method.
Introduction and Motivation
Human gaze estimation is incorporated into many different Human-Robot Interaction applications because of the importance of the gaze as a social non-verbal cue for interaction It drives many different social cognitive mechanisms such as joint attention, intention prediction, and task coordination and provides an explainable behaviour for others Affective states are also represented in the gaze behaviour The ability to perceive and understand the social cues affects the effectiveness and efficiency of the whole interaction experience Achieving high accuracy in gaze estimation is a key enabler to reach a seamless Human-Robot interaction task This paper presents an applied evaluation for the latest gaze estimation methods in a standard HRI scenario, specifically when the human and the robot are engaged in a shared task space (e.g., table surface).
Methodology and Setup
The study considers a setup where a human interacts with a humanoid robot in a shared working space with a dominant plane (e.g. a table). The system uses stereo-vision, face detection and gaze estimation methods to estimate the gaze point in the plane
** HW setup: The NICO robotic platform is equipped with two See3CAM CU135 wide field-of-view cameras inside the robot’s head, and the robot’s torso is placed on a table with a built-in display **
** Camera Calibration: Camera calibration is performed with a checkerboard pattern using [29], and knowledge of camera intrinsics allows images to be rectified by removing distortion **
** Gaze Estimation Pipeline: To estimate human gaze, the system first detects bounding boxes of faces in images from both robots’ cameras taken simultaneously, then triangulates the center points to obtain a 3D point. The normalized gaze direction is obtained from a gaze-estimation network, and this ray is transformed into the coordinate system of the work surface plane π. The point of intersection with the plane π in Cπ is found by solving for α = −x′3/d′3. **
Evaluation Dataset and Metrics
The evaluation dataset was collected with six participants, five men and one woman in the age range 20-25. The participants were instructed to look at numbered white squares of the grid shown on the display in the workspace shared with the robot NICO. In total, our evaluation dataset contains 315 images from each camera.
** Evaluation Metrics: To evaluate gaze accuracy, we use the angle between the ground truth gaze direction and the gaze estimated by the network to report mean angular error. We also calculate the error in terms of distance on a shared work surface by finding where the gaze ray intersects the work surface plane and calculating its distance from the center of the target square. Since distance can grow asymptotically with increasing angular gaze error, we report median distance instead of mean. We also report Precision@Xcm for points estimated within 10, 20, and 50 cm distance error. **
Evaluation Results
Four recent methods were included in the evaluation: GazeTR [9], L2CS [10], 3DGazeNet [11], and Gaze3D [12].
** Performance Comparison: Table 1 shows the results of the evaluated methods for the full dataset and its subsets with participants either wearing glasses or not wearing glasses. L2CS and 3DGazeNet perform the best in terms of angular error. **
** Distance Error Analysis: In terms of the distance in the plane defined by the shared work surface, 3DGazeNet performs the best with a median distance of 16.48 cm. This suggests that localization of gaze within a shared work surface is limited in its accuracy. **
** Error Distributions: Figure 3 shows the cumulative distribution of the angular and distance errors, indicating that GazeTR reach mean angular error similar to the error they achieve on the Gaze360 dataset [13] (∼ 10◦). **
Discussion and Limitations
Current off-the-shelf gaze estimation methods can be beneficial in HRI systems, but their limitations must be considered. When used in a scenario similar to ours, the gaze information can only provide an estimate with a resolution of tens of centimeters. This accuracy can inform the potential layout of the shared space such that this accuracy is sufficient to discriminate between objects or relevant portions of the shared working area. Conversely, gaze estimation networks could be used to provide only broader cues, such as whether the person is looking at the robot or at the shared workspace. The inaccuracy could be accounted for and supplemented with information from other modalities in a multimodal perception system or by leveraging a broader context of the task. Performing multiple estimates using several images (video frames) could also lead to greater accuracy in aggregate or real-time information about the certainty of prediction accuracy. As future work, our dataset could be extended to include a greater variety of participants, additional robotic platforms, or different scenarios. Additionally, the evaluation could be strengthened by also considering videos providing additional temporal context.
Conclusion
In order to perform the evaluation, we have collected an annotated dataset using the robotic platform NICO. We found that the methods perform within the expected accuracy in terms of the mean angular error. However, when the estimated gaze is used to determine a specific gaze point in a planar working space shared by the robot and the human, the error within the plane expressed in terms of metric distance remains relatively high, with a 16 cm median error for the best performing method. Based on our findings, we provide several recommendations on how to incorporate off-the-shelf gaze estimation methods in HRI systems. The evaluation code is available at https://github.com/kocurvik/nico gaze. The dataset will be made available upon acceptance.
Data Availability Statement
The dataset will be made available upon acceptance. The evaluation code is available on https://github.com/kocurvik/nico gaze.
Improvements for AI systems
-
A multimodal perception system can leverage gaze estimation to
provide only broader cues (e.g., whether the person is looking at the robot or at the shared workspace), which could still be useful,
supplementing its inherent inaccuracy with context from other modalities or task knowledge. -
Gaze-informed spatial planning can be enhanced by using the estimated gaze point on a planar working space to
inform the potential layout of the shared space, such that this accuracy is sufficient to discriminate between objects or relevant portions of the shared working area,
allowing for more context-aware object placement or workspace organization. -
Temporal gaze prediction models can be developed by
Performing multiple estimates using several images (video frames) could also lead to greater accuracy in aggregate or real-time information about the certainty of prediction accuracy.
This would enable HRI systems to produce more robust, real-time gaze predictions by aggregating data across a short video sequence.
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models