Estimating Accurate Hand Pose in Camera Space with Vision Transformer
cs.CV, cs.AI, cs.GR
Submitted: 2026-09-21
Updated: 2026-09-22
Code: https://github.com/Mine268/CS-ViT
License: http://creativecommons.org/licenses/by/4.0/
The gist: Monocular RGB-based hand pose estimation has emerged as a critical research frontier in computer vision.
Terminology
Abstract
Monocular RGB-based hand pose estimation has emerged as a critical research frontier in computer vision. The local hand pose estimation methods predict hand poses relative to the wrist, while global hand pose estimation also requires estimating the wrist's position in the camera coordinate system. However, this camera-space estimation confronts two fundamental challenges: (1) depth ambiguity in monocular settings, and (2) the coupling effect of hand local poses and global wrist positions in the perspective projections. In particular, this coupling reflects that the projections are jointly determined by local hand poses, wrist positions, and camera intrinsics. To overcome these challenges, our framework proposes two key innovations: Transformation-Isomorphism Supervision for hand-depth information extraction and Perspective Information Embedding for resolving above coupling effect of local pose and wrist position, both integrated within the mainstream encoder-decoder architecture. Besides, we propose a novel framerate-aware multi-dataset training strategy for sequential pose refinement. Our fully integrated approach achieves at most 37.1% superiority in CS-MJE over SOTA on HO3D. Project page: https://github.com/Mine268/CS-ViT.
Sources
- A Simple Framework for Contrastive Learning of Visual Representations
- Masked Autoencoders Are Scalable Vision Learners
- Scaling Laws for Neural Language Models
- Pre-Training for 3D Hand Pose Estimation with Contrastive Learning on Large-Scale Hand Images in the Wild
- Decoupled Weight Decay Regularization
- MediaPipe: A Framework for Building Perception Pipelines
- Bringing Inputs to Shared Domains for 3D Interacting Hands Recovery in the Wild
- Recovering 3D Hand Mesh Sequence from a Single Blurry Image: A New Dataset and Temporal Unfolding
- WiLoR: End-to-end 3D Hand Localization and Reconstruction in-the-wild
- PeCLR: Self-Supervised 3D Hand Pose Estimation from monocular RGB via Equivariant Contrastive Learning
- RoFormer: Enhanced Transformer with Rotary Position Embedding
- ViTPose++: Vision Transformer for Generic Body Pose Estimation
- Dyn-HaMR: Recovering 4D Interacting Hand Motion from a Dynamic Camera
- HaWoR: World-Space Hand Motion Reconstruction from Egocentric Videos
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models