CLoSeR: Closing the Loop for Long-Context Streaming Reconstruction
cs.CV, cs.RO
Submitted: 2026-10-01
Updated: 2026-10-01
Code: https://github.com/MoyangLi00/CLoSeR
Terminology
Sources
- Geometric Context Transformer for Streaming 3D Reconstruction
- Easi3R: Estimating Disentangled Motion from DUSt3R Without Training
- TTT3R: 3D Reconstruction as Test-Time Training
- VGGT-Long: Chunk it, Loop it, Align it -- Pushing VGGT's Limits on Kilometer-scale Long RGB Sequences
- D2-Net: A Trainable CNN for Joint Detection and Description of Local Features
- VGG-T$^3$: Offline Feed-Forward 3D Reconstruction at Scale
- ZipMap: Linear-Time Stateful 3D Reconstruction via Test-Time Training
- OKVIS2: Realtime Scalable Visual-Inertial SLAM with Loop Closure
- SlideSLAM: Sparse, Lightweight, Decentralized Metric-Semantic SLAM for Multi-Robot Navigation
- VGGT-SLAM 2.0: Real-time Dense Feed-forward Scene Reconstruction
- VGGT-SLAM: Dense RGB SLAM Optimized on the SL(4) Manifold
- DINOv2: Learning Robust Visual Features without Supervision
- FastVGGT: Training-Free Acceleration of Visual Geometry Transformer
- Learning to (Learn at Test Time): RNNs with Expressive Hidden States
- DeepV2D: Video to Depth with Differentiable Structure from Motion
- AMB3R: Accurate Feed-forward Metric-scale 3D Reconstruction with Backend
- $\pi^3$: Permutation-Equivariant Visual Geometry Learning
- InfiniteVGGT: Visual Geometry Grounded Transformer for Endless Streams
- MonST3R: A Simple Approach for Estimating Geometry in the Presence of Motion
- LoGeR: Long-Context Geometric Reconstruction with Hybrid Memory
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models