RegVGGT: Sustainable Visual Geometry Grounding for Streaming via Regulated Memory
cs.CV
Submitted: 2026-09-20
Updated: 2026-09-20
Code: https://github.com/amao996/RegVGGT
Terminology
Sources
- TTT3R: 3D Reconstruction as Test-Time Training
- ScanNet: Richly-annotated 3D Reconstructions of Indoor Scenes
- FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- $\textit{S}^3$Gaussian: Self-Supervised Street Gaussians for Autonomous Driving
- MapAnything: Universal Feed-Forward Metric 3D Reconstruction
- STream3R: Scalable Sequential 3D Reconstruction with Causal Transformer
- WinT3R: Window-Based Streaming Reconstruction with Camera Token Pool
- Robust Incremental Structure-from-Motion with Hybrid Features
- Evict3R: Training-Free Token Eviction for Memory-Bounded Streaming Visual Geometry Transformers
- DINOv2: Learning Robust Visual Features without Supervision
- FastVGGT: Training-Free Acceleration of Visual Geometry Transformer
- XStreamVGGT: Extremely Memory-Efficient Streaming Vision Geometry Grounded Transformer with KV Cache Compression
- Attention Is All You Need
- 3D Reconstruction with Spatial Memory
- $\pi^3$: Permutation-Equivariant Visual Geometry Learning
- Point3R: Streaming 3D Reconstruction with Explicit Spatial Pointer Memory
- EmbodiedOcc: Embodied 3D Occupancy Prediction for Vision-based Online Scene Understanding
- InfiniteVGGT: Visual Geometry Grounded Transformer for Endless Streams
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models