eVGGT: An Efficient Geometry-Aware Vision Encoder for Visuomotor Policies

arXiv:2509.15880 · cs.RO · Submitted 2025-09-19 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: Today's paper: "eVGGT: An Efficient Geometry-Aware Vision Encoder for Visuomotor Policies".

Dev: Existing RGB-based imitation learning approaches typically employ traditional vision encoders such as ResNet or ViT, which lack explicit 3D reasoning capabilities.

Rosa: First, who's behind it and why it matters.

Paper summary: Dev: The core thesis here is that while models like VGGT offer robust spatial understanding, they are too costly for practical robotic deployment because of their high computational expense. So the authors propose eVGGT as a solution by distilling the knowledge from VGGT into a much lighter model.

Rosa: It’s about making these geometry-grounded models accessible, and they claim that this distillation process results in eVGGT being nearly nine times faster and five times smaller than the original VGGT while keeping its strong three dee reasoning abilities intact.

Taro: That size reduction is significant for real-world deployment because it means we can run these complex vision models on hardware that isn't a super high-end workstation.

Rosa: And they don't just stop there; they show how simple the integration is, stating that you can replace traditional 2D vision encoders’ latent space with the latent space of their proposed geometry-aware encoder in standard imitation learning baselines.

Conclusion: Rosa: Looking at the work on "eVGGT: An Efficient Geometry-Aware Vision Encoder for Visuomotor Policies," it seems the main point is taking a powerful three dee vision understanding model and making it practical for actual robots by making it much smaller and faster without losing its core geometric insight.

Dev: The authors are essentially showing that you don't have to sacrifice strong spatial awareness just because the computational cost is too high for real-time operation on a robot platform.

Taro: This has big implications because if we can deploy these three dee reasoning capabilities more efficiently, it opens up possibilities for robots to handle much more complex and unstructured environments than they could before.

Rosa: It really boils down to bridging the gap between high-end research models and usable robotic systems by focusing on efficiency while maintaining that crucial geometry awareness.

Dev: And when we consider the performance gains mentioned, like that six point five percent improvement in success rate over standard encoders in manipulation tasks, it suggests that this efficiency boost isn't just about running faster; it translates into better actual task performance for the AI policies themselves.

Rosa: That’s what’s exciting; it means we get both speed and accuracy improvements simultaneously when we introduce this geometry-aware encoding into frameworks like ACT or DP.

Taro: I wonder how this efficiency plays out when things go wrong in the physical world, Rosa? If the three dee understanding is implicit, what happens if the environment presents a scenario that falls outside that implicit understanding?

Dev: That’s a critical point for me; we need to know where these limitations lie when we move from simulation to real-world testing.

Rosa: That's exactly what I want to explore next, and it leads us into how this system holds up outside the controlled lab setting.

Dinh Vuong, Minh Nhat Vu, Ian Reid

Department of Computer Vision, Mohammed bin Zayed University of Artificial Intelligence

cs.RO

Submitted: 2025-09-19

Updated: 2026-09-28

Comments: IROS 2026. Project page: https://evggt.github.io/

License: http://creativecommons.org/licenses/by-nc-sa/4.0/

Importance score: 82/100

The gist: Existing RGB-based imitation learning approaches typically employ traditional vision encoders such as ResNet or ViT, which lack explicit 3D reasoning capabilities.

Key concepts

Geometry-Aware Vision Encoder
This is a specialized encoder that processes visual data not just as 2D pixels but also by understanding the underlying 3D structure or geometry of the scene. It allows the model to 'see' spatial relationships, which is crucial for tasks like robotic manipulation where understanding object shapes and positions is necessary.
Knowledge Distillation
This technique is used to create eVGGT. A large, powerful 'teacher' model (VGGT) trains a smaller 'student' model (eVGGT) by having the student mimic the outputs of the teacher. This transfers the complex geometric understanding from the large model into a much more efficient and faster version.
Imitation Learning Integration
This involves taking eVGGT's 3D spatial features and plugging them directly into existing imitation learning algorithms, such as ACT or DP. Instead of training new vision components, the policy heads learn to interpret the geometry-aware features provided by eVGGT for action prediction.

Terminology

Summary

Existing RGB-based imitation learning approaches typically employ traditional vision encoders such as ResNet or ViT, which lack explicit 3D reasoning capabilities. This work investigates the integration of geometry-aware visual representations into robotic manipulation, demonstrating that incorporating a geometry-aware vision encoder into imitation learning frameworks yields up to 6.5% improvement over standard vision encoders in success rate across single- and bi-manual manipulation tasks in both simulation and real-world settings.

The gist: eVGGT is a lightweight geometry-aware encoder distilled from VGGT, which, when integrated into imitation learning frameworks like ACT and DP, improves performance by up to 6.5% over standard vision encoders while being nearly 9× faster and 5× smaller than the original VGGT.

How it works

The core of the proposed method involves creating eVGGT, an efficient geometry-aware encoder distilled from VGGT [1]. This distillation process is achieved through a knowledge distillation framework that trains the student model (eVGGT) on the outputs of a pretrained teacher model (VGGT). The training loss is defined as:

Ldistill = X N i=1∥θ std i −θ tch i∥ 2+∥Dstd i −Dtch i∥ 2+∥P std i −P tch i∥ 2 (Equation 1)

To address noise in depth maps, a gradient loss is incorporated:

Lgrad = X N i=1∥∇xDstd i −∇xDtch i∥1+∥∇yDstd i −∇yDtch i∥1 (Equation 2)

The student model, eVGGT, is designed to be lightweight by reducing the number of transformer blocks from L = 24 in VGGT to L = 4. Furthermore, it utilizes a smaller backbone: we replace the DINOv2-ViT-L backbone with the smaller DINOv2-ViT-S. This results in eVGGT being nearly 9× faster and 5× smaller than VGGT while preserving strong geometric reasoning capabilities.

Integration into Robotic Policies

The distilled encoder is integrated into imitation learning by replacing traditional 2D vision encoders with the latent space of eVGGT. The approach is simple yet effective: we replace traditional 2D vision encoders’ latent space with that of our proposed geometry-aware encoder in standard IL baselines. This allows for direct use of 3D spatial representations without additional task-specific fine-tuning, as the policy heads handle the projection of these features.

This integration is demonstrated across two representative imitation learning approaches:

  1. ACT (Action Chunking with Transformer): eVGGT produces a latent feature yt = eVGGT(ot) which is incorporated into both the transformer encoder and decoder to generate predicted action chunks.

  2. DP (Diffusion Policy): The observation Ot is encoded into a global-context embedding Yt = eVGGT(Ot), which provides informative 3D guidance for denoising in the reverse process of the diffusion model, conditioning the denoising network fω(A(k)t, Yt, k).

Evaluation and Performance Gains

Experiments across simulation and real-world settings demonstrate significant performance improvements. In RoboTwin simulation, ACT+eVGGT achieved improvements of 2.1% over the original ACT and DP+eVGGT achieved 6.5% over the original DP. On ManiSkill, DP+eVGGT reached a success rate of 0.96 for Pick Cube, outperforming RDT and DP with RGBD input by 1.7% and 4.3%, respectively.

When compared against other vision methods in Table III, eVGGT outperforms DP+ResNet by over 5% and DP+ViT by 11.5%. Notably, the paper observes that DP3’s performance drops by 20.5% when using eVGGT point clouds, highlighting that the implicit encoding design is more effective as it integrates scene geometry without requiring object-level segmentation.

Efficiency and Real-World Transfer

A key contribution is the efficiency of eVGGT, which considerably reduces model size by 5× and 2× compared to VGGT and achieves at least 9× faster inference. This speedup makes extensive IL experiments feasible. Furthermore, real-world evaluations using a Kinova Gen3 robot show that DP+eVGGT achieves a 72% success rate vs. 54% for the baseline in the Push Cube task, confirming that the geometry-aware encoder successfully transfers from simulation to the real world.

Limitations and Broader Impact

The paper notes limitations, stating that its point clouds are dense and unstructured, which can reduce performance in 3D-based imitation learning baselines, and that its 3D understanding remains implicit. The authors suggest exploring future work on "

Improvements for AI systems

Here are specific improvements that can be made to existing AI systems based on the proposed eVGGT and its integration into Imitation Learning (IL) frameworks:

  1. Improve manipulation success rates by integrating a geometry-aware vision encoder (eVGGT) into Imitation Learning policies (ACT or DP). Specifically, this yields up to a 6.5% improvement over standard RGB encoders in both single- and bi-manual manipulation tasks across simulation and real-world settings.

  2. Enhance the ability of policies to handle complex geometric reasoning by leveraging eVGGT's latent space instead of traditional 2D vision encoders, leading to more robust learned behaviors for object interaction and spatial alignment.

  3. Achieve significant efficiency gains in deployment by utilizing the distilled eVGGT encoder, which is nearly 9x faster and 5x smaller than the original VGGT model while preserving strong 3D reasoning capabilities, making it practical for real-time robotic systems with hardware constraints.

  4. Enable high-fidelity 3D scene reconstruction (depth maps and point clouds) during policy learning by using eVGGT to extract geometric features, allowing policies to condition their actions on explicit 3D spatial context derived from the visual input.

  5. Improve generalization of learned manipulation skills by incorporating a gradient loss function (Eq. 2) during knowledge distillation, which helps mitigate student model overfitting and ensures the predicted depth maps exhibit better smoothness and consistency compared to standard distillation losses.

  6. Improve robustness in 3D reconstruction under challenging conditions by applying data augmentation techniques (color jitter, grayscale conversion, Gaussian blur) specifically to the student input images during distillation training, which is shown to improve 3D reconstruction accuracy (e.g., reducing Chamfer distance gap).

  7. Achieve state-of-the-art performance in robotic manipulation tasks by utilizing DP+eVGGT policies on real robots, demonstrating that geometry-aware representations can support success rates comparable to large-scale pretrained methods like RDT and approaches using explicit depth modalities (DP3).

  8. Enable effective knowledge transfer from simulation to the real world for robotic manipulation by training policies with eVGGT on diverse synthetic demonstrations and then deploying them on real robots, resulting in measurable improvements in success rates (e.g., 72% success rate for Push Cube vs. 54% baseline).

  9. Develop more accurate camera pose estimation within imitation learning frameworks by using eVGGT's features as input to the transformer encoder component of models like ACT, leading to better temporal consistency in action chunk prediction.

Abstract

Geometry-grounded vision models, such as VGGT, have emerged as robust visual encoders, providing essential geometric priors for robotic manipulation. However, the high computational cost of these models often leads to slow inference, limiting their practical applications in real-world robotics. This paper introduces eVGGT, a lightweight geometry-aware vision encoder distilled from the high-performing VGGT. Our findings demonstrate two primary advantages: i) integrating eVGGT into imitation learning frameworks (including ACT and Diffusion Policy) yields up to a 6.3% improvement in success rate over standard 2D encoders across bimanual and single-arm tasks in both simulation and real-world settings with variable viewpoints; ii) eVGGT achieves a nearly 5 times speedup and a 63% reduction in memory usage compared to state-of-the-art geometry-aware encoders while maintaining comparable task performance. These results suggest that eVGGT substantially alleviates the performance-latency bottleneck that has limited geometry-aware visuomotor policies in real-world deployment.

Sources

Related papers