InfiniHand: Streaming World-Space Hand Motion Estimation from Egocentric Video
cs.CV
Submitted: 2026-09-28
Updated: 2026-09-28
Project page: https://infinihand.github.io
Terminology
Sources
- Geometric Context Transformer for Streaming 3D Reconstruction
- Hamba: Single-view 3D Hand Reconstruction with Graph-guided Bi-Scanning Mamba
- EgoDex: Learning Dexterous Manipulation from Large-Scale Egocentric Video
- Metric3Dv2: A Versatile Monocular Geometric Foundation Model for Zero-shot Metric Depth and Surface Normal Estimation
- Sapiens2
- EgoWAM: World Action Models Beyond Pixels with In-the-Wild Egocentric Human Data
- EgoLive: A Large-Scale Egocentric Dataset from Real-World Human Tasks
- OmniHands: Towards Robust 4D Hand Mesh Recovery via A Versatile Transformer
- VGGT-SLAM: Dense RGB SLAM Optimized on the SL(4) Manifold
- Bringing Inputs to Shared Domains for 3D Interacting Hands Recovery in the Wild
- MASt3R-SLAM: Real-Time Dense SLAM with 3D Reconstruction Priors
- WiLoR: End-to-end 3D Hand Localization and Reconstruction in-the-wild
- 3D Hand Pose Estimation in Everyday Egocentric Images
- EgoVerse: An Egocentric Human Dataset for Robot Learning from Around the World
- M^3: Dense Matching Meets Multi-View Foundation Models for Monocular Gaussian Splatting SLAM
- The Surprising Effectiveness of Video Diffusion Models for Hand Motion Reconstruction
- Metric3D: Towards Zero-shot Metric 3D Prediction from A Single Image
- HaWoR: World-Space Hand Motion Reconstruction from Egocentric Videos
- LoGeR: Long-Context Geometric Reconstruction with Hybrid Memory
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models