SGAP-Gaze: Scene Grid Attention Based Point-of-Gaze Estimation Network for Driver Gaze
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "SGAP-Gaze: Scene Grid Attention Based Point-of-Gaze Estimation Network for Driver Gaze".
Jane: Driver gaze estimation is essential for understanding situational awareness, and this paper proposes SGAP-Gaze,
Tom: First, who's behind it and why it matters.
Title and authors: Tom: So Jane, we're diving into this new paper today about SGAP-Gaze. It sounds like they're tackling a really important problem in driver monitoring by using scene context to help figure out where a driver is looking.
Jane: It definitely seems like they are trying to make the gaze estimation more reliable by not just looking at the face, but also what's happening in the traffic scene around them. It’s about adding that extra layer of information, isn't it?
Lu: Exactly! This whole idea of explicitly incorporating scene information into gaze modeling through a Scene Grid Attention mechanism is really interesting because it moves beyond just analyzing the driver's face in isolation, which is where many current systems might struggle to keep up with real-world complexity.
Meng: From an engineering standpoint, I’m curious how they actually manage all that input. They are combining facial features, eye movements, and scene context; that sounds like a massive data fusion challenge for any system to handle efficiently in real time.
Lalam: It seems like the core of this work is integrating these different modalities—the face, the eye features, and the scene grid tokens—into a single unified gaze descriptor called "zgaze" before they even start predicting anything. That kind of cross-modal understanding is where AI can really start to get contextually aware.
Tom: Right, so that unified vector is what feeds into their prediction heads for both the three dee direction and the Point-of-Gaze estimation, which is what we're after in driver monitoring systems.
Jane: And the results they are reporting are quite solid; they achieved a mean pixel error of one hundred four point seven three on the UD-FSG dataset and sixty-three point four eight on the LBW dataset, which they say is a twenty-three point five percent reduction compared to previous top models.
Lu: That reduction in error is significant when you consider how much context the scene information brings to the table; it shows that this approach actually yields tangible improvements over existing state-of-the-art methods for driver gaze estimation.
Meng: A twenty-three point five percent improvement on those specific metrics is a good benchmark, but I wonder about the practical deployment aspect—how robust is this system when the traffic environment gets extremely dense or unpredictable?
Lalam: The paper addresses that by using scene features projected into a sequence of grid-based tokens, which allows the attention mechanism to weigh relevant parts of the scene grid based on what’s in that unified gaze intent vector. That sounds like it directly improves spatial consistency by focusing computational power where it matters most for gaze prediction.
Tom: So, what they are doing is taking those scene features and using them as a query within a Transformer-based attention mechanism to refine the initial facial and eye embeddings before making the final point estimation. It’s a layered approach, isn't it?
Jane: Precisely; they start by extracting features from the face—using something like a YOLOv8 Face-Eye-Iris detector—and then they use those extracted features alongside scene tokens to compute attention scores for predicting the Point-of-Gaze.
Title and authors: Lu: The methodology involves several distinct steps, starting with facial geometry detection and then moving into multi-stream feature extraction, where they handle face features, weighted eye features emphasizing the iris location with a spatial Gaussian weighting function, and finally the scene feature extraction into those grid tokens.
Meng: I see how that feeds into their fusion module; taking those different streams—face embeddings, eye embeddings, and scene representations—and compressing them into that "zgaze" vector before applying attention sounds like a very structured way to ensure all relevant inputs are accounted for simultaneously.
Lalam: And after getting that unified gaze intent vector, they use it as the query to compute attention over the scene grid, which then generates the final Point-of-Gaze prediction, including a vertical residual correction term to handle discretization bias. That attention mechanism is key to their robustness.
Tom: So if I'm tracking this correctly, they are using that scene grid attention not just for general saliency but specifically to guide where the point of gaze should be located on the actual scene layout, which is a big step forward.
Jane: And when they look at the three dee Gaze Direction Estimation, they use a hybrid loss function combining cosine and Euclidean losses to penalize angular deviation while also using an L2 regression term for coordinate accuracy.
Lu: That hybrid loss setup suggests they are balancing the need for accurate directionality with precise location prediction, which is smart because both are crucial for understanding driver intent in three-dimensional space.
Meng: I'm thinking about the implications here; if we can get a more accurate three dee gaze direction vector, it means we move beyond just knowing where a driver is looking on a 2D screen and start understanding their actual line of sight in the physical world.
Lalam: That capability could really enhance proactive safety systems by allowing the AI to anticipate potential hazards based on where the driver's attention is directed in three dimensions, not just what's happening right in front of them.
Tom: It definitely suggests that these systems could evolve from simply flagging distractions to providing a much deeper assessment of driving behavior based on actual visual focus.
Jane: And we have to remember that this SGAP-Gaze model is designed for the UD-FSG dataset, which is characterized by heterogeneous traffic environments in Kanpur city, so its performance there speaks directly to real-world applicability.
Lu: The fact that they successfully incorporated the scene images into the gaze modeling framework demonstrates a strong path toward developing more contextually aware AI agents that can operate effectively across varied and complex driving scenarios.
Meng: From a practical standpoint, this means we are building systems that can handle real-world variability better because they aren't stuck relying on just one type of input, which is where most simpler models fall apart when conditions change.
Lalam: And from an AI culture perspective, this level of multi-modal integration shows how powerful it is when different pieces of information—visual, spatial, and contextual—are woven together by a strong attention mechanism to produce a coherent result.
Title and authors: Tom: So to wrap up this section on the SGAP-Gaze paper, we've seen how they combine facial geometry, eye features with spatial weighting, and scene grid tokens through an attention mechanism to predict both three dee direction and the Point-of-Gaze.
Jane: It really shows that by explicitly modeling gaze as a function of both facial features and scene information, they can achieve a measurable improvement in accuracy over existing methods on real driving data.
Lu: This work paves the way for future research where we can explore even more complex attention patterns, perhaps looking at how to adapt this grid attention mechanism to different types of visual information beyond just traffic scenes.
Meng: For implementation, the next hurdle will be making sure that this entire pipeline runs fast enough on embedded systems without sacrificing that level of contextual awareness they achieved.
Lalam: The potential for future work is huge; thinking about how we can use this framework to integrate gaze estimation into larger multi-agent collaboration systems, perhaps allowing agents to coordinate their attention based on shared scene understanding.
Tom: What a compelling discussion on SGAP-Gaze; it’s clear that by enriching the input data with scene context via a Scene Grid Attention mechanism, they've made a meaningful contribution to driver monitoring accuracy.
Jane: It certainly gives us more confidence that AI systems can move toward providing more nuanced safety assessments by understanding where a driver is truly directing their attention in the environment.
Lu: This paper really shows that when you systematically integrate different data streams through carefully designed fusion modules, you get a much richer representation of the underlying phenomenon being studied.
Meng: So while the results are encouraging on benchmarks, I’m eager to see how this architecture performs when we introduce noise or occlusions that are common in actual driving conditions.
Lalam: The implications for culture are that we start seeing AI not just as a tool for detection, but as a system capable of understanding situational awareness across multiple sensory inputs simultaneously.
Tom: That’s what makes these discussions so valuable; it’s not just about the numbers, it's about how this technology can actually inform better safety decisions on the road.
Jane: Absolutely; this research moves us closer to systems that understand the context of a driver's actions, which is a big step toward truly intelligent monitoring tools.
Lu: It’s exciting to see how researchers build these complex attention-based models that manage such diverse input types effectively for tasks like gaze estimation.
Meng: I’m looking forward to seeing what engineering constraints they overcome next when trying to deploy this on less powerful hardware.
Lalam: And I'm excited about the long-term vision of using this kind of cross-modal understanding to build more sophisticated, contextually aware AI agents in physical spaces.
Tom: Alright team, that’s our look at SGAP-Gaze; a really solid piece of research that adds significant depth to driver gaze estimation. We’ll be back after the break with another interesting paper on arXiv!
The paper's summary: Tom: So we've been looking at the heavy lifting in SGAP-Gaze, and now it's time to break down exactly what this paper is claiming about its performance and overall goal for driver monitoring.
Jane: That makes sense, Tom; we need to get past the technical jargon and understand what this model actually does for us in terms of real-world accuracy.
Lu: The core summary highlights that SGAP-Gaze achieves a mean pixel error of one hundred four point seven three on the UD-FSG dataset and sixty-three point four eight on the LBW dataset, which they claim is a twenty-three point five percent reduction compared to previous top models for driver gaze estimation.
Meng: A twenty-three point five percent reduction in mean pixel error is substantial when you consider how much context we're talking about; that suggests a real lift in reliability for those kinds of monitoring systems on the road.
Lalam: What's really important from the summary is their approach, which is modeling gaze as a function of both facial features and scene features using a Scene Grid Attention mechanism.
Tom: Exactly, so they aren't just looking at the face in isolation anymore; they are explicitly incorporating what’s happening in the traffic scene into how they model where the driver is looking.
Jane: In simple terms, this means the system gains situational awareness because it doesn't just see a driver's eyes; it sees what that driver is paying attention to in their surroundings, like a nearby object or road sign.
Lu: That explicit incorporation of scene information through that grid-based attention is what separates this work from simpler models and gives them the edge on complex driving scenarios.
Meng: From an engineering standpoint, the fact they’ve managed to fuse those different streams—the face data, the weighted eye features, and those scene tokens—into one unified gaze descriptor is a big technical feat.
Lalam: The implication for culture here is huge because it moves us toward AI that truly understands context; instead of just flagging in-cabin distractions, we could build systems that monitor the driver’s actual attention flow across their entire environment.
Tom: That's right; this isn't just about better numbers on a dataset; it’s about creating a monitoring system that offers a much deeper assessment of driving behavior.
Jane: It gives us more confidence that these AI tools can provide early alerts for potential inattentive driving behaviors by mapping exactly where the driver is directing their attention in the environment.
Lu: And since they achieved better performance on heterogeneous datasets like UD-FSG, we can expect this to translate well into real-world deployment across various traffic conditions, not just a controlled lab setting.
Meng: I'm thinking about how this could change how we design these systems; instead of relying on fixed gaze points, we might be able to predict dynamic attention shifts based on scene context.
Lalam: This level of contextual awareness could fundamentally alter safety assessments by providing a much richer picture of a driver's state in real time.
Tom: So, what this paper really summarizes is that by weaving facial cues and scene context together through this specific attention mechanism, they've built a more accurate way to pinpoint where a driver is looking in the world.
Jane: And it’s exciting because it shows how multi-modal data fusion can lead to tangible improvements in safety applications.
Lu: It certainly lays the groundwork for future work where we can explore even more complex attention patterns that might incorporate other types of environmental information beyond just traffic scenes.
Meng: I'm curious, though, about the practical deployment on less powerful hardware; if this level of accuracy requires a lot of processing power, that's a hurdle we need to clear.
Lalam: The long-term potential is that we start seeing AI not just as a tool for detection but as a system capable of understanding situational awareness across multiple sensory inputs simultaneously in physical spaces.
Tom: That’s the big picture, folks; this paper shows us exactly how to enrich the input data with scene context via a Scene Grid Attention mechanism to predict both three dee direction and Point-of-Gaze.
The paper's improvements: Tom: So we've looked at how SGAP-Gaze works and its current performance numbers, and now we’re going to focus on what the authors suggest they could improve moving forward.
Jane: That’s a smart way to look at it, Tom; seeing where they think the next step is helps us understand the roadmap for making these systems even better.
Lu: The paper outlines four specific improvements: utilizing a Scene Grid Attention Mechanism for gaze localization, integrating multi-modal feature fusion across face, eye, iris, and scene context, enhancing robustness to peripheral gaze targets in outer regions, and improving spatial consistency via attention weighting.
Meng: Those points sound promising for real-world deployment because they specifically address the limitations we talked about earlier regarding peripheral regions where existing models often struggle.
Lalam: The fusion aspect is key; by combining those different data streams—face features, eye movements weighted by iris position, and scene context—they aim to create a much more comprehensive understanding of the driver's focus.
Tom: So it’s not just about one model getting better; it’s about building a system that intelligently combines different types of visual information for a richer output.
Jane: And that leads directly to the prediction heads, where they use that unified gaze vector to generate both a three dee gaze direction and the Point-of-Gaze estimation using attention over the scene grid.
Lu: The improvement in spatial consistency through attention weighting is what I find most interesting; it means the model can dynamically decide which parts of the driving scene are most relevant for estimating that final point of gaze.
Meng: If that mechanism successfully guides where the prediction focuses, it suggests a way to make the system more adaptive to fast-changing environments on the road.
Lalam: The impact here is profound because when AI systems can consistently localize gaze across all spatial ranges, it allows us to build driver monitoring systems that provide nuanced safety assessments for everything from in-cabin distraction to external hazards.
Tom: It sounds like the authors are focused on moving from a static estimation to a more dynamic, context-aware gaze prediction system.
Jane: Precisely; they’re trying to ensure the gaze estimation is not just accurate in the center of the face, but also robust when that gaze drifts toward less obvious or peripheral areas.
Lu: This work also points toward future research where we can look at how this grid attention mechanism could be adapted to handle different types of visual information that aren't strictly traffic scenes.
Meng: From an engineering standpoint, the next challenge will be implementing these complex fusion and attention layers in a way that keeps the latency low enough for real-time use on embedded systems.
Lalam: The vision here is that we move toward AI agents capable of understanding driver intent across their entire environment, which could fundamentally alter how we design safety protocols for autonomous vehicles and human drivers alike.
Tom: So, while they’ve given us a solid architecture and good results, the paper’s improvements show they are actively thinking about making this system more robust and adaptable in complex driving situations.
Conclusion: Tom: So we’ve wrapped up our deep dive into SGAP-Gaze, and to recap, this paper introduces a Scene Grid Attention based Point-of-Gaze estimation network that uses scene context to model where a driver is looking.
Jane: That's right; it’s all about making gaze estimation more robust by integrating the visual information of the traffic scene directly into the model.
Lu: The main conclusion is that this method achieves better localization accuracy than state-of-the-art models by explicitly modeling gaze as a function of both facial geometry and environmental features.
Meng: It seems like a solid piece of work, but I still have to ask about the computational cost; how scalable is this architecture for deployment in high-speed or very complex environments?
Lalam: The cultural impact here is that it pushes AI toward systems that possess genuine situational awareness, allowing us to build monitoring tools that understand the driver's actual focus in a rich, contextual manner.
Tom: That’s the big picture we’re talking about; this research moves us closer to monitoring systems that provide a much deeper assessment of driving behavior than just simple distraction detection.
Jane: It definitely gives us more confidence that AI can provide those nuanced safety assessments by understanding where a driver is truly directing their attention in the environment.
Lu: I think the future work will be really interesting, especially exploring how this grid attention mechanism could be adapted to handle different kinds of visual information beyond just traffic scenes.
Meng: I’m still focused on the engineering side; if we can optimize that multi-modal fusion module for lower latency, then it becomes something we can actually put into a production monitoring system.
Lalam: The real advancement is in how this type of cross-modal understanding improves safety protocols across the board, because systems that understand context are inherently safer systems.
Tom: Well said; this paper on SGAP-Gaze really shows us how systematically integrating different data streams through a Scene Grid Attention mechanism can lead to measurable improvements in accuracy.
Jane: It’s clear that by making gaze estimation dependent on the scene, they’ve created a model that's more grounded in reality, which is what we need for reliable monitoring.
Lu: This work lays excellent groundwork for exploring more complex attention patterns later on that might incorporate even more diverse environmental inputs.
Meng: We just need to keep pushing those engineering constraints so this powerful concept can actually run efficiently on the hardware we use every day.
Lalam: Ultimately, SGAP-Gaze demonstrates how AI can evolve into a system capable of understanding situational awareness across multiple sensory inputs simultaneously in physical spaces.
Pavan Kumar Sharma, Pranamesh Chakraborty
cs.CV
Submitted: 2026-04-21
Updated: 2026-09-28
Code: https://github.com/pavans20/Urban-Driving-Face-Scene-Gaze-Dataset
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 74/100
The gist: Driver gaze estimation is essential for understanding situational awareness, and this paper proposes SGAP-Gaze, a Scene Grid Attention based Point-of-Gaze estimation network that explicitly
Key concepts
- Scene Grid Attention
- This technique allows the model to selectively pay attention to specific spatial locations within the traffic scene image. It treats the scene as a grid of tokens, enabling the network to weigh which parts of the environment are most important when determining where a driver is looking.
- Multi-Modal Feature Fusion
- The model combines different types of information—facial features, eye features, and scene features—into one unified representation. This fusion process creates a comprehensive understanding of the gaze by merging cues from both the person's face and their surroundings.
- Point-of-Gaze (PoG) Estimation
- This is the core task: predicting the exact 3D location in space where a driver's eyes are directed. The model predicts this by using an attention mechanism to find the most relevant parts of the scene grid, resulting in a precise estimate of gaze direction.
- Gaze = f(F, S)
- This mathematical formulation describes the core idea: that the driver's gaze location is a function of two main inputs: F (Facial features) and S (Scene features). The model learns to map these two distinct sets of data into a single, accurate gaze prediction.
Terminology
Summary
Driver gaze estimation is essential for understanding situational awareness, and this paper proposes SGAP-Gaze, a Scene Grid Attention based Point-of-Gaze estimation network that explicitly incorporates scene information into gaze modeling to achieve robust driver PoG estimation.
The gist
The proposed SGAP-Gaze model achieves a mean pixel error of 104.73 on the UD-FSG dataset and 63.48 on the LBW dataset, achieving a 23.5% reduction in mean pixel error compared to state-of-the-art driver gaze estimation models.
Dataset and Data Collection
The study introduces the Urban Driving-Face Scene Gaze (UD-FSG) dataset, which comprises synchronized driver face and traffic scene images.
This dataset was collected using an instrumented vehicle equipped with two USB webcams synchronized via a G-Streamer application, while driver gaze ground truth was obtained using a Pupil Invisible eye tracker. The data involved 35 male professional drivers and included 3,73,488 driver face and scene image pairs along with 2D gaze ground truth coordinates.
The dataset is characterized by its heterogeneous traffic environments,
incorporating variations in traffic density and lighting conditions across different time of day in Kanpur city.
Model Architecture
The proposed SGAP-Gaze architecture consists of four major components: (1) Facial geometry detection module, (2) Multi-stream feature extraction module, (3) Multi modal feature fusion module, and (4) Gaze prediction head. The problem is formulated as modeling gaze location as a function of facial features and scene features: gaze = f(F, S).
(1) Facial Geometry Extraction Module:
The face image is processed using a custom YOLOv8 Face-Eye-Iris (FEI) detector, which achieved an mAP of 95.7%. This module detects the face, eye, and iris regions,
and it leverages physiological principles to infer the spatial coordinates of undetected irises using a validity-based gating mechanism
when reliable iris coordinates are unavailable.
(2) Multi-Stream Feature Extraction Module:
This module processes different inputs through separate streams:
-
Face Feature Extraction: A pretrained ResNet18 backbone is used to extract hierarchical facial representations, resulting in four global facial feature vectors, each of 256-D.
-
Gaussian-Weighted Eye Feature Extraction: Eye features are extracted using a ResNet-18 backbone on cropped eye images, followed by resizing and padding to meet the required input resolution (3×224×224). Crucially, a
spatial Gaussian weighting function
centered at the iris location is applied to emphasize the iris region, yielding aGaussian-weighted global embedding.
-
Scene Feature Extraction: A ResNet18 backbone extracts spatial features from the scene image, which are then projected into a sequence of grid-based tokens (7×7 feature map), resulting in 49 tokens, where each token corresponds to a spatial scene grid.
(3) Multi-Modal Feature Fusion Module:
The modality-specific embeddings are combined to form a joint gaze descriptor:
-
Facial Features Representation: Face features, eye features (eL and eR), and iris centers are concatenated into a unified face embedding, which is then fused with the eye embeddings to create a representation capturing
fine-grained eye-driven gaze cues.
-
Gaze Intent Representation: These modality-specific embeddings are compressed into a unified gaze intent vector, denoted as
zgaze,
which serves as the query for attention computation over the scene grid.
Gaze Prediction Head
The model employs two prediction heads:
-
3D Gaze Direction Estimation: The 3D gaze direction vector is regressed directly from the fused gaze embedding using a fully connected layer, with the resulting vector normalized to unit length.
-
Attention based Point of Gaze Prediction (PoG): This involves fusing the face and scene grid. The
zgaze
vector acts as the query, while scene features are projected into key vectors. Attention scores are computed via dot-product similarity:αi = exp(z gaze i T k i).
The final PoG is obtained as the weighted average of the predefined grid center coordinates:ˆp = X N i=1 αi.
A vertical residual correction term, ∆p, is applied to compensate for discretization bias before yielding the final prediction.
Loss Functions and Evaluation
The model utilizes specific loss functions for supervision:
(3D Gaze Direction Loss):
A hybrid cosine and Euclidean loss is used: Ldir = λ1 Lcos + λ2 LL2,
where Lcos penalizes angular deviation, and LL2 is an L2 regression term.
Improvements for AI systems
As a fastidious and diligent researcher, I have analyzed the proposed SGAP-Gaze framework and its contributions to driver gaze estimation. Based on this research, here are specific improvements for existing AI systems:
The proposed SGAP-Gaze framework can be leveraged to create significantly more robust and context-aware driver monitoring systems by integrating multi-modal scene understanding directly into gaze localization.
Here are the specific improvements and what the improved AI system can do:
-
Utilization of a Scene Grid Attention Mechanism for Gaze Localization:
-
Integration of Multi-Modal Feature Fusion (Face, Eye, Iris, and Scene Context):
-
Enhanced Robustness to Peripheral Gaze Targets (Outer Regions):
-
Improved Spatial Consistency via Attention Weighting:
The improved AI system can achieve the following specific capabilities:
-
A driver monitoring system can accurately predict the exact spatial location of a driver's gaze in the environment (Point-of-Gaze or PoG) by fusing facial cues, eye movements (weighted by iris position), and real-time scene context simultaneously.
-
The system will provide a 3D gaze direction vector, allowing for precise understanding of the driver's intended line of sight in three-dimensional space, which is superior to models relying solely on face features.
-
The model will exhibit superior localization accuracy compared to state-of-the-art methods (like GazePTR) by achieving a 23.5% reduction in mean pixel error on complex, real-world driving datasets (UD-FSG).
-
The system will demonstrate enhanced robustness when the driver looks toward the edges or corners of the scene (peripheral regions), where existing models typically fail due to reliance on unreliable eye tracking cues, leading to higher accuracy across all spatial ranges.
-
The improved system can be deployed in advanced Driver Monitoring Systems (DMS) to provide early alerts for potential inattentive driving behaviors by precisely mapping where the driver's attention is directed, offering a more nuanced safety assessment than current systems that only detect in-cabin distractions.
Abstract
Driver gaze estimation is essential for understanding the driver's situational awareness of surrounding traffic. Existing gaze estimation models use driver facial information to predict the Point-of-Gaze (PoG) or the 3D gaze direction vector. We propose a benchmark dataset, Urban Driving-Face Scene Gaze (UD-FSG), comprising synchronized driver-face and traffic-scene images. The scene images provide cues about surrounding traffic, which can help improve the gaze estimation model, along with the face images. We propose SGAP-Gaze, Scene-Grid Attention based Point-of-Gaze estimation network, trained and tested on our UD-FSG dataset, which explicitly incorporates the scene images into the gaze estimation modelling. The gaze estimation network integrates driver face, eye, iris, and scene contextual information. First, the extracted features from facial modalities are fused to form a gaze intent vector. Then, attention scores are computed over the spatial scene grid using a Transformer-based attention mechanism fusing face and scene image features to obtain the PoG. The proposed SGAP-Gaze model achieves a mean pixel error of 104.73 on the UD-FSG dataset and 63.48 on LBW dataset, achieving a 23.5% reduction in mean pixel error compared to state-of-the-art driver gaze estimation models. The spatial pixel distribution analysis shows that SGAP-Gaze consistently achieves lower mean pixel error than existing methods across all spatial ranges, including the outer regions of the scene, which are rare but critical for understanding driver attention. These results highlight the effectiveness of integrating multi-modal gaze cues with scene-aware attention for a robust driver PoG estimation model in real-world driving environments.
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models