Hand-4DGS: Feed-Forward 3D Gaussian Splatting for 4D Hand Reconstruction from Egocentric Videos
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Hand-4DGS: Feed-Forward 3D Gaussian Splatting for 4D Hand Reconstruction from Egocentric Videos".
Jane: Dynamic 3D hand reconstruction from egocentric videos is essential for next-generation computing platforms such as AR/VR and AI glasses, and this paper introduces Hand-4DGS,
Tom: First, who's behind it and why it matters.
Title and authors: Tom: Welcome back everyone. We're diving into a really interesting piece today, a paper titled "Hand-4DGS: Feed-Forward three dee Gaussian Splatting for 4D Hand Reconstruction from Egocentric Videos <ref:2606.19156#pg0,Feed-Forward 3D Gaussian Splatting for 4D Hand Reconstruction from Egocentric Videos>." Jane, we've been hearing about how crucial it is to get dynamic hand models working in AR and AI glasses, so what should we focus on when we look at this?
Jane: I think the title itself tells us a lot, Tom; it’s about using a feed-forward approach with three dee Gaussian Splatting to reconstruct 4D hands directly from egocentric videos <ref:2606.19156#pg0>. That means instead of doing all that heavy optimization work, the system just takes the video and spits out the dynamic hand geometry and motion pretty quickly.
Lu: From a theoretical standpoint, this tackles a major hurdle in using three dee Gaussian Splatting, which usually demands significant time for optimization, by proposing a feed-forward way to predict those parameters directly from input images <ref:2606.19156#pg0>. It bypasses that traditional bottleneck entirely.
Meng: And the fact that it targets egocentric videos instead of needing multiple fixed views is huge for practical deployment because it removes the need for bulky camera setups or complex multi-view tracking systems right away.
Lalam: I see this paper as a significant cultural step forward because if we can generate highly realistic, dynamic hand models from just a single video feed, it opens up possibilities for immersive interactions that were previously too computationally expensive to achieve consistently.
Tom: Exactly! So what’s the core idea behind Hand-4DGS <ref:2606.19156#pg0>? Jane, can you explain the main mechanism in simple terms for our listeners?
Jane: Certainly. The paper explains that instead of traditional methods, Hand-4DGS uses an encoder to pull features from each frame, and then decoder networks map those features straight into the parameters that define a three dee Gaussian—things like position, rotation, scale, opacity, and color <ref:2606.19156#pg0>.
Lu: It’s a feed-forward three dee Gaussian Splatting approach where the network predicts these parameters directly from the input images rather than iteratively optimizing them over time <ref:2606.19156#pg0,feed-forward 3D Gaussian Splatting>. That’s what allows for that fast inference speed they claim is around sixty frames per second <ref:2606.19156#pg0>.
Meng: That speed is critical for real-time applications, which is where I focus. If it can run at sixty frames per second without needing lengthy optimization steps, that drastically cuts down on the latency we have to worry about when interacting with a user or controlling a robot remotely.
Lalam: From my perspective as the language model, this fast reconstruction capability means we can integrate these hand models into more interactive AI agents where hand gestures and movements need to be interpreted instantly rather than being delayed by complex rendering pipelines.
Tom: That’s the speed advantage right there. Jane, how does it handle the structure of a hand, which is inherently articulated? Is that something they addressed specifically?
Title and authors: Jane: Yes, they introduced a mesh-guided representation to leverage the consistent articulated structure of hands. They first predict a MANO mesh using features from the encoder, and then they add positional embeddings onto that surface by interpolating from vertex embeddings.
Lu: The paper details how these vertex embeddings act as conditional latents for predicting individual Gaussian parameters, which is a clever way to link the surface geometry directly to the three dee representation <ref:2606.19156#pg0>. Furthermore, they sample additional Gaussians within each triangular face using barycentric interpolation of those three neighboring vertex embeddings for higher resolution details.
Meng: So it’s not just a smooth blob of points; it’s building a structured mesh first and then refining the appearance on that structure, which makes sense from an engineering standpoint for maintaining structural integrity during motion.
Lalam: The idea of using mesh structure as a prior to guide the Gaussian prediction is powerful; it ensures that even when the hand moves rapidly, the underlying skeletal structure remains coherent in the reconstruction.
Tom: That sounds like a very solid way to ensure fidelity. Jane, what about keeping things stable when dealing with video sequences where motion can look jittery?
Jane: To stabilize those features and reduce temporal jitter, Hand-4DGS incorporates temporal convolution layers <ref:2606.19156#pg0>. Instead of looking at each frame in isolation, the framework aggregates features from consecutive frames within a temporal window of size 2k plus one.
Lu: This results in temporally consistent features called f' in R times N times D', where the dimension D' is smaller than the original dimension D, which helps stabilize the feature representation before it goes into the MANO and embedding predictions.
Meng: Reducing that dimensionality by using temporal context sounds like a smart way to manage computational load while still capturing dynamic hand motion, which is a real win for deployment feasibility.
Lalam: That temporal awareness is crucial for our AI systems; it means the model isn't just reacting to the current frame but understands the trajectory of the hand movement across several frames, which makes its predictions much more intuitive.
Tom: So we’ve covered how it works and how they stabilize things temporally. Now, let's talk about what they did to train this system and what kind of supervision they used in "Hand-4DGS: Feed-Forward three dee Gaussian Splatting for 4D Hand Reconstruction from Egocentric Videos <ref:2606.19156#pg0,Feed-Forward 3D Gaussian Splatting for 4D Hand Reconstruction from Egocentric Videos>."
Jane: The training uses a multi-stage strategy. They start by using pseudo ground truth derived from an off-the-shelf pose estimator to supervise the predicted mesh vertices.
Lu: They gradually reduce this reliance on that initial pose supervision to encourage the model to rely more on image supervision later in the process. Then, they combine L1 loss and Structural Similarity Index Measure or D-SSIM between what the model renders and the actual ground truth images.
Title and authors: Meng: Combining those losses is standard practice, but using both helps them enforce both pixel accuracy and structural coherence simultaneously, which is a robust training signal for complex geometry like hands.
Lalam: I find that combining image loss with structural similarity encourages the model to focus on preserving the overall shape and appearance of the hand, not just hitting individual pixels perfectly. It’s about learning the right visual grammar of a hand.
Tom: That makes sense, it balances precise detail with overall structure. Jane, what are they saying about how this method performs when tested on real-world challenges?
Jane: They evaluated Hand-4DGS on challenging egocentric datasets like H2O and ARCTIC, showing significant improvements over baselines in both 4D hand reconstruction and hand pose estimation accuracy <ref:2606.19156#pg0>.
Lu: The results show that their approach significantly outperforms the existing baselines, which include methods like HUGS and EVA, in both reconstructing the 4D hand geometry and estimating the actual hand pose from those videos <ref:2606.19156#pg0>.
Meng: When I look at the results on ARCTIC, it shows stability across different scenes and challenging interactions, which speaks directly to its potential for deployment in varied real-world environments without needing extensive per-scene fine-tuning.
Lalam: The generalization aspect is really what stands out; they show that the model trained on one set of videos can perform feed-forward inference on completely unseen test sequences while maintaining stable geometry across those new cases.
Tom: That generalization capability is a major point, Jane. It means we don't need to re-optimize the entire scene every time we want to see a new video; we just run the feed-forward network. So, what’s the final word on what this paper achieves?
Jane: In conclusion, Hand-4DGS is presented as the first feed-forward framework capable of reconstructing dynamic 4D hands directly from egocentric videos <ref:2606.19156#pg0>. Its innovations include that mesh-guided representation and temporal convolution layers for consistency, and it achieves fast inference at around sixty frames per second without needing ground truth annotations during the inference step <ref:2606.19156#pg0>.
Lu: The main contributions boil down to the feed-forward nature, the mesh guidance with vertex embeddings, and using temporal convolution to produce temporally consistent features that reduce jittering. It’s a complete pipeline designed for speed and generalization.
Meng: From an engineering standpoint, achieving that sixty FPS inference while maintaining high quality is what makes this practical; it moves the technology from research papers into something deployable in consumer hardware like AI glasses <ref:2606.19156#pg0>.
Lalam: This paper moves the needle because it allows us to build more responsive and realistic interactive systems where hand dynamics are central, which really enhances the user experience in any immersive application we develop.
Tom: So, that’s a wrap on Hand-4DGS; it’s a very clever way to tackle 4D reconstruction from egocentric video by focusing on fast inference and strong generalization <ref:2606.19156#pg0>. We’ll keep an eye on how this feeds into our next set of papers.
The paper's summary: Tom: So, to wrap up what we just covered, Hand-4DGS is essentially taking video input and using a feed-forward network to directly spit out a full 4D hand—meaning its shape and its movement over time—without getting bogged down in long optimization routines that usually plague this kind of three dee reconstruction work <ref:2606.19156#pg0>.
Jane: That’s right, Tom; the core innovation is bypassing that heavy iterative optimization process by having the network predict all those complex parameters like position, rotation, scale, opacity, and color all at once from the raw video data. It makes generating these dynamic models much faster than what we usually see in existing research.
Lu: From a creative standpoint, I think the way they use that mesh-guided representation is fascinating; it’s not just throwing points around randomly; they are first building a consistent skeleton or MANO shape and then decorating it with Gaussian details, which really grounds the reconstruction in actual human anatomy.
Meng: I’m focused on how fast this actually runs on hardware. The fact that the system claims sixty frames per second inference is what gets my attention; if we can deploy this kind of reconstruction capability in real-time AR glasses, it opens up a whole new level of interaction for our robotic systems.
Lalam: I see the biggest cultural impact here being how this capability enhances our ability to create realistic digital representations; when we can generate highly plausible, moving hand models from simple video input, it really advances the fidelity of any virtual avatar or interactive agent we build.
Tom: Exactly! And the generalization aspect they highlighted is huge; they’re showing that a model trained on one set of videos can actually perform well on completely new, unseen sequences without needing to be retrained for every single scene.
Jane: That ability to generalize means we don't need massive amounts of specific ground-truth data for every single video we want to process, which simplifies the pipeline immensely for real-world use cases.
Lu: It suggests that the learned representations within this framework are capturing really fundamental, high-level dynamics of human hand motion across different contexts rather than just memorizing a few specific poses.
Meng: If we can reduce our dependency on those expensive per-scene annotations during inference, that lowers the barrier for deploying advanced perception systems in less controlled or more dynamic environments.
Lalam: This advancement contributes to a culture where creating sophisticated, interactive AI agents becomes much more accessible because the underlying visualization tools are becoming faster and smarter.
Tom: Speaking of which, this paper lays a really solid foundation for using three dee Gaussian Splatting in a way that’s practical for live video streams; it’s not just theoretical anymore. Where do we go from here, and what are the next steps for this technology?
The paper's improvements: Tom: So, we've talked about how Hand-4DGS works under the hood and what they’ve already achieved in terms of speed and reconstruction quality; now let's look at what they suggest as improvements for future versions of the framework <ref:2606.19156#pg0>.
Jane: That’s right, Tom; the authors are pointing out that while it handles dynamic motion well, there are still areas where they can refine the temporal modeling to get even smoother results in very fast or chaotic movements.
Lu: I think their suggestion to further integrate different types of temporal aggregation might be really interesting because it could allow the model to distinguish between slow, deliberate hand movements and very quick, jittery gestures more effectively.
Meng: From an engineering standpoint, if they can refine the feature extraction part—that encoder component—to be more robust against occlusions in egocentric footage, that would significantly increase its practical utility for robotics applications.
Lalam: I think these suggested improvements really touch on how we can make AI models better at understanding human intent; refining those temporal features means the model is learning a richer grammar of hand motion.
Tom: That’s a great point, Lu; improving the feature extraction robustness is key because that’s where the input information gets filtered before it even hits the three dee reconstruction part.
Jane: And for us listeners, what this implies is that we can expect future versions of these models to be even more reliable in capturing the subtle nuances of human gesture, which could eventually lead to much more natural-feeling virtual interactions.
Lu: I think it also points toward exploring how different learning regimes—the paper touched on those earlier—could be combined to optimize the temporal component specifically for different types of hand dynamics.
Meng: If they can manage that complexity without making the feed-forward inference slower than sixty frames per second, that's when we see a real engineering win for deployment on edge devices.
Lalam: For our AI culture, this suggests a future where AI avatars don't just move smoothly, but their hand movements reflect true intent and context in a way that feels deeply human.
Tom: So it sounds like the next phase involves fine-tuning the temporal aggregation to handle extreme motion better while keeping inference snappy. Jane, what about limitations? Where does this approach still fall short?
Conclusion: Tom: So, to wrap up our discussion on Hand-4DGS: Feed-Forward three dee Gaussian Splatting for 4D Hand Reconstruction from Egocentric Videos, we’ve seen how this framework achieves fast, high-quality 4D hand reconstruction directly from egocentric video using a feed-forward approach <ref:2606.19156#pg0,Gaussian Splatting for 4D Hand Reconstruction from Egocentric Videos>.
Jane: That's right, Tom; it’s a really clever way to bypass the slow optimization steps that usually take up so much time in traditional three dee Gaussian Splatting methods by predicting the geometry and motion parameters straight from the input frames.
Lu: The real power here, as I see it, is how they integrate those mesh embeddings with the Gaussian parameters; it creates a structured representation that holds together even when the hand is moving through complex poses.
Meng: From an engineering standpoint, what this means for us is that we have a much more efficient path to building perception systems for robotics and AR because we aren't stuck waiting hours for reconstruction before we can take action.
Lalam: For me, the most significant implication is how this moves us toward a culture where digital interactions feel incredibly natural because the hand models are dynamic and responsive in real-time.
Tom: And I think that generalization capability they showed—the model working on unseen videos—is what really makes this practical for deployment rather than just a neat research paper.
Jane: Absolutely, Tom; it means we can create avatars or AR overlays that adapt to any new video stream without needing a completely new training session every single time.
Lu: That suggests we could see incredibly rich, diverse interactions in future AI applications where the environment is constantly changing and unpredictable.
Meng: I'm still thinking about the limitations they mentioned; they noted that while it’s fast, achieving perfect fidelity under extremely high-speed or highly occluded scenarios might still require some fine-tuning beyond what a pure feed-forward setup can handle perfectly.
Lalam: And even with those caveats, the potential for improving the way we visualize and interact with AI is huge; imagining hand gestures that are instantly understood by an AI agent is a powerful vision for our future.
Tom: That's right; so Hand-4DGS gives us a much faster, more generalizable tool to tackle 4D hand reconstruction in real-world scenarios <ref:2606.19156#pg0>.
Jane: It’s certainly an exciting piece of research that really pushes the boundaries of how we model dynamic human motion.
Lu: We definitely need to keep watching how they expand on those temporal modeling suggestions because that could unlock even more complex motion understanding.
Meng: I’m eager to see what the next iteration looks like from a deployment feasibility standpoint, especially regarding latency targets.
Lalam: This paper shows us that AI can build representations of the physical world—like hands—that are fast enough and flexible enough to truly enhance our digital experience.
Yonsei University · Electronics and Telecommunications Research Institute (ETRI) · Microsoft Spatial AI Lab
cs.CV
Submitted: 2026-06-17
Updated: 2026-10-07
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
Importance score: 92/100
The gist: Dynamic 3D hand reconstruction from egocentric videos is essential for next-generation computing platforms such as AR/VR and AI glasses, and this paper introduces Hand-4DGS, the first feed-forward
Key concepts
- 3D Gaussian Splatting
- This is a technique used to represent 3D scenes or objects by using many small, translucent 3D shapes called Gaussians. Instead of complex meshes, the method uses these splats to model the hand's surface. It allows for high-quality rendering and fast inference because it directly predicts visual properties like position and color from input images.
- Mesh-Guided Representation
- The paper uses a mesh (MANO) to define the overall shape of the hand. Positional embeddings are placed on this mesh surface using information derived from vertex embeddings. These vertex embeddings act as conditional inputs, helping the model accurately predict the specific 3D parameters for each individual Gaussian point on the hand.
- Temporal Convolution Layers
- These layers process features across consecutive video frames to create 'temporally consistent features.' Instead of treating each frame in isolation, they aggregate information from a small window of frames. This helps stabilize the model's predictions and reduces unwanted flickering or jitter when reconstructing fast hand movements in a video sequence.
- Feed-Forward Framework
- This means the system works by directly mapping input data (video frames) to output parameters (3D Gaussian properties) without needing lengthy, iterative optimization steps. This design is key to achieving very fast inference speeds, such as 60 frames per second, making it suitable for real-time applications like AR/VR.
Terminology
Summary
Dynamic 3D hand reconstruction from egocentric videos is essential for next-generation computing platforms such as AR/VR and AI glasses, and this paper introduces Hand-4DGS, the first feed-forward framework for reconstructing dynamic 4D hands directly from egocentric videos, enabling fast (∼60 FPS) inference and strong generalization.
How it works
The core of Hand-4DGS is a feed-forward 3D Gaussian Splatting approach that bypasses the heavy optimization time typically required by traditional methods by predicting 3D Gaussian parameters directly from input images. The framework consists of an encoder, Enc(·), which extracts per-frame features ft = Enc(It) from RGB frames, and decoder networks (lightweight MLPs) that map these features to Gaussian parameters such as position (µ), rotation (r), scale (s), opacity (σ), and color (c). This allows for fast inference
and generalization to unseen videos without requiring a pose estimator during inference.
Mesh-Guided Representation for Hands
To leverage the consistent articulated structure of hands, the method introduces a mesh-guided representation. First, it predicts a MANO mesh using the encoder features f. Then, positional embeddings are introduced on these surfaces by barycentric interpolation from vertex embeddings (Equation 1). These vertex embeddings serve as conditional latents for predicting individual Gaussian parameters.
To capture higher resolution details, additional Gaussians are sampled within each triangular face of the mesh, and their embeddings are obtained through barycentric interpolation of the three neighboring vertex embeddings (Equation 2), producing Gaussian embeddings at a higher resolution that smoothly vary across the hand surface.
Temporal Modeling and Feature Enhancement
To stabilize feature representations and reduce temporal jitter inherent in video sequences, Hand-4DGS incorporates temporal convolution layers. Instead of using per-frame features independently, the framework aggregates features from consecutive frames within a temporal window of size 2k + 1 (Equation 5), resulting in "temporally consistent features ft' ∈ R×N×D′ with reduced dimension D′ < D. These
temporal-aware features are then used for subsequent MANO and embedding predictions, enabling the model to
capture dynamic hand motion in egocentric videos."
Training Objectives and Supervision
The training process employs a multi-stage supervision strategy. Vertex supervision is applied first using pseudo ground truth obtained from an off-the-shelf pose estimator: we supervise the predicted mesh vertices using pseudo ground truth obtained from an off-the-shelf pose estimator
(Equation 6). This supervision is gradually reduced to encourage reliance on image supervision. Image supervision combines L1 loss and Structural Similarity Index Measure (D-SSIM) between the rendered and ground truth images: we combine L1 loss and structural dissimilarity (D-SSIM) between the rendered and ground truth images
(Equation 7). Regularization is applied to constrain Gaussian scales within a range [smin, smax] and encourage opacity values close to 1. The total loss combines these components: The complete loss combines all components: L = λvert L vert + λimg L img + L reg
(Equation 9).
Experimental Validation and Generalization
Hand-4DGS was evaluated on challenging egocentric datasets like H2O and ARCTIC. The framework demonstrates significant improvements over baselines in both 4D hand reconstruction and hand pose estimation.
Crucially, the method shows strong generalization capabilities: while Hand-4DGS (Per-Video) is optimized directly on a target video, the Hand-4DGS (Generalized) variant trained on multiple sequences can perform feed-forward inference on unseen test sequences,
maintaining stable geometry across all test cases
and achieving quality comparable to the per-scene model. Furthermore, 2D image supervision is shown to effectively improve hand pose estimation by encouraging the model to learn reconstructions that are better aligned with the input sequences.
The method also achieves superior pose accuracy compared to initial estimates provided by HaMeR.
Key Contributions
The paper introduces Hand-4DGS as the first feed-forward framework capable of reconstructing dynamic 4D hands directly from egocentric videos.
Its key innovations include:
-
A mesh-guided representation that predicts MANO shape and uses vertex embeddings for conditional Gaussian attributes.
-
The incorporation of temporal convolution layers to produce
temporally consistent features
that reduce jittering. -
The ability to achieve fast inference (∼60 FPS) and strong generalization without requiring expensive 3D hand pose ground-truth annotations during inference.
-
A mechanism where 2D image supervision refines both predicted mesh vertices and Gaussian parameters during training, improving reconstruction fidelity and pose accuracy.
Improvements for AI systems
As a fastidious researcher, I have analyzed the core contributions of Hand-4DGS and identified several high-impact areas where this framework can be leveraged to significantly improve existing AI systems.
Here are the specific improvements and capabilities derived from this research:
-
Improved Efficiency in Real-Time AR/VR Applications
-
Enhanced Generalization for Unseen Environments
-
Robustness to Complex Egocentric Dynamics (Occlusions and Fast Motion)
-
Reduced Dependency on Expensive Ground-Truth Annotations
-
High-Fidelity, Low-Latency 4D Hand Reconstruction in AI Glasses:
This system can reconstruct dynamic, articulated hand geometry (shape and appearance) from a single, moving camera view at high frame rates (e.g., 60 FPS). Unlike methods requiring multi-view inputs or per-video optimization, this framework enables real-time rendering of plausible 4D hand poses directly from egocentric video streams.
- Autonomous Robotic Manipulation and Teleoperation:
The system can serve as a visual perception layer for robotic systems (e.g., teleoperation or dexterous manipulation). By providing accurate 3D hand geometry and pose estimates in real-time, robots can better plan interactions, grasp objects with precision, and execute bimanual tasks by understanding the dynamics of human hands in an egocentric setting.
- Next-Generation AI/VR Avatar Synthesis:
This framework allows for the creation of highly realistic digital avatars that exhibit complex hand motions. The ability to generalize to unseen videos means a single model can be trained on diverse data and immediately generate plausible, dynamic hand animations for novel scenarios without needing scene-specific retraining.
- Augmented Reality (AR) Interaction Fidelity:
In AR applications, this system can provide highly realistic overlays of virtual hands interacting with the user's real environment or other digital objects. The mesh-guided representation ensures that these virtual hands maintain structural integrity and realistic appearance even when viewed from novel angles or under occlusions, significantly increasing the immersion factor.
- Improved Hand Pose Estimation Accuracy in Unconstrained Settings:
By utilizing 2D image supervision (Gaussian Splatting loss) alongside structural priors (MANO mesh embeddings), the system can refine initial hand pose estimates derived from external pose estimators (like HaMeR). This allows for more accurate tracking of hand joints in uncalibrated or complex egocentric videos, leading to better control and interaction in AI agents.
- Reduced Training Overhead via Feed-Forward Architecture:
The feed-forward paradigm eliminates the need for per-scene optimization, drastically reducing the computational time required for inference on new data compared to traditional 3D Gaussian Splatting methods that require extensive scene-specific optimization. This makes deploying robust hand reconstruction models faster and more cost-effective.
Abstract
Dynamic 3D hand reconstruction from egocentric videos is essential for next-generation computing platforms such as AR/VR and AI glasses. Despite its importance, most prior works focus either on multi-view 3D hand reconstruction or on 4D human body reconstruction. Egocentric 4D hand reconstruction remains difficult due to rapid hand and camera motion, hand-object and inter-hand interactions, and inherent ambiguity from single-view observations. To address these challenges, we introduce Hand-4DGS, a feed-forward framework for dynamic 4D hand reconstruction from egocentric videos. Our approach incorporates a mesh-guided representation for structural priors and temporal convolutions to model dynamic motion. We evaluate our framework on H2O and ARCTIC, two egocentric hand-object interaction datasets, and show improvements over baselines. Our model can efficiently adapt to unseen videos from datasets that are not included in training. In addition, image supervision through differentiable Gaussian rasterization provides an additional signal for appearance optimization and pose adjustment during training and test-time optimization, without ground-truth 3D hand pose annotations.
Sources
- MVSplat: Efficient 3D Gaussian Splatting from Sparse Multi-View Images
- HandSCS: Structural Coordinate Space for Animatable Hand Gaussian Splatting
- Forge4D: Feed-Forward 4D Human Reconstruction and Interpolation from Uncalibrated Sparse-view Videos
- EVA-Gaussian: 3D Gaussian-based Real-time Human Novel View Synthesis under Diverse Multi-view Camera Settings
- BIGS: Bimanual Category-agnostic Interaction Reconstruction from Monocular Videos via 3D Gaussian Splatting
- Adversarial Motion Modelling helps Semi-supervised Hand Pose Estimation
- JGHand: Joint-Driven Animatable Hand Avater via 3D Gaussian Splatting
- Predicting 4D Hand Trajectory from Monocular Videos
- Dyn-HaMR: Recovering 4D Interacting Hand Motion from a Dynamic Camera
- GUAVA: Generalizable Upper Body 3D Gaussian Avatar
- HaWoR: World-Space Hand Motion Reconstruction from Egocentric Videos
- IDOL: Instant Photorealistic 3D Human Creation from a Single Image
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models