OTT3R: Multi-View 3D Reconstruction and Fast Dataset Generation at 1% Compute
cs.CV
Submitted: 2026-09-28
Updated: 2026-09-28
Code: https://github.com/TheFourthKaramazov/OTT3R
Terminology
Sources
- Multi-View 3D Reconstruction using Knowledge Distillation
- Distilling the Knowledge in a Neural Network
- Training Compute-Optimal Large Language Models
- Scaling Laws for Neural Language Models
- Distill3R: A Pipeline for Democratizing 3D Foundation Models on Commodity Hardware
- Towards Spatio-Temporal World Scene Graph Generation from Monocular Videos
- DINOv3
- Understanding Multi-View Transformers
- InfiniteVGGT: Visual Geometry Grounded Transformer for Endless Streams
- Attention Itself Could Retrieve.RetrieveVGGT: Training-Free Long Context Streaming 3D Reconstruction via Query-Key Similarity Retrieval
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models