WATCH: World-aware Allied Trajectory and pose reConstruction for Camera and Human
cs.CV
Submitted: 2025-09-04
Updated: 2026-09-22
Terminology
Sources
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- TokenHMR: Advancing Human Mesh Recovery with a Tokenized Pose Representation
- HumanPlus: Humanoid Shadowing and Imitation from Humans
- Humans in 4D: Reconstructing and Tracking Humans with Transformers
- PACE: Human and Camera Motion Estimation from in-the-wild Videos
- GENMO: A GENeralist Model for Human MOtion
- Joint Optimization for 4D Human-Scene Reconstruction in the Wild
- DINOv2: Learning Robust Visual Features without Supervision
- CameraHMR: Aligning People with Perspective
- WHAM: Reconstructing World-grounded Humans with Accurate 3D Motion
- Deep Patch Visual Odometry
- VGGT: Visual Geometry Grounded Transformer
- DUSt3R: Geometric 3D Vision Made Easy
- PromptHMR: Promptable Human Mesh Recovery
- TRAM: Global Trajectory and Motion of 3D Humans from in-the-wild Videos
- Decoupling Human and Camera Motion from Videos in the Wild
- WHAC: World-grounded Humans and Cameras
- HumanMM: Global Human Motion Recovery from Multi-shot Videos
- Synergistic Global-space Camera and Human Reconstruction from Videos
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models