Skel-WAM: A Hand-Skeleton-Conditioned World Action Model for Human-to-Robot Manipulation Transfer
cs.RO
Submitted: 2026-09-18
Updated: 2026-09-18
Project page: https://healorcai.github.io/Skel-WAM
Terminology
Sources
- EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos
- Scalable Vision-Language-Action Model Pretraining for Robotic Manipulation with Real-Life Human Activity Videos
- H2R: A Human-to-Robot Data Augmentation for Robot Pre-training from Videos
- EgoEngine: From Egocentric Human Videos to High-Fidelity Dexterous Robot Demonstrations
- Ego2Robot: Scalable Robot Data Synthesis from Egocentric Human Data
- HandEdit: A Unified Benchmark for Egocentric Human-to-Robot Dexterous Hand Image Editing
- EgoScale: Scaling Dexterous Manipulation with Diverse Egocentric Human Data
- Emergence of Human to Robot Transfer in Vision-Language-Action Models
- Ego-Pi: VLA Fine-Tuning for Ego-Centric Human and Robot Data
- EgoSteer: An Open-Source Full-Stack System Towards Steerable Dexterous Manipulation from Egocentric Videos
- HumanEgo: Zero-Shot Robot Learning from Minutes of Human Egocentric Videos
- Dexterity from Smart Lenses: Multi-Fingered Robot Manipulation with In-the-Wild Human Demonstrations
- Robots Acquire Manipulation Skills in Seconds from a Single Human Video
- FlowWAM: Optical Flow as a Unified Action Representation for World Action Models
- MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models
- Point Policy: Unifying Observations and Actions with Key Points for Robot Manipulation
- World Models for Learning Dexterous Hand-Object Interactions from Human Videos
- OSCAR: Omni-Embodiment Action-Conditioned World Model for Robotics
- GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving