Weave: Learning Whole-Body Dexterous Loco-Manipulation from Human-Object Interactions
cs.RO, cs.AI, cs.LG
Submitted: 2026-09-15
Updated: 2026-09-15
Comments: 10 pages, 5 figures. Project website: https://xiaohu-art.github.io/Weave/
Code: https://github.com/kevinzakka/mink
Project page: https://xiaohu-art.github.io/Weave
License: http://creativecommons.org/licenses/by/4.0/
The gist: Learning humanoid-object interaction requires coordinating whole-body balance, locomotion, and dexterous hand contact to control both robot and object motion.
Terminology
Abstract
Learning humanoid-object interaction requires coordinating whole-body balance, locomotion, and dexterous hand contact to control both robot and object motion. Human demonstrations provide examples of coordinated interaction, but transferring these behaviors to humanoid robots requires learning how to establish and maintain effective contacts under different embodiments and dynamics. We present Weave, a unified framework for learning whole-body dexterous humanoid-object interaction from captured human demonstrations. Weave first converts captured human-object interactions into executable robot-object references through contact-aware retargeting and approach-motion completion. At its core is a contact- and geometry-aware policy that jointly commands 29 body joints and 12 actuated finger joints across multiple objects and interaction sequences. Evaluation across nine objects yields a 92.5% success rate on trained interactions and, without any additional training, 65.0% on sequences never seen during training. We additionally release 9,000 physically executed rollouts spanning 23 hours, providing robot-object trajectories with contact annotations for downstream interaction-policy learning and physically consistent HOI motion generation. Project website: https://xiaohu-art.github.io/Weave/
Sources
- HDMI: Learning Interactive Humanoid Whole-Body Control from Human Videos
- OmniRetarget: Interaction-Preserving Data Generation for Humanoid Whole-Body Loco-Manipulation and Scene Interaction
- ResMimic: From General Motion Tracking to Humanoid Whole-body Loco-Manipulation via Residual Learning
- SceneBot: Contact-Prompted General Humanoid Whole Body Tracking with Scene-Interaction
- Emergent Active Perception and Dexterity of Simulated Humanoids from Visual Reinforcement Learning
- InterPrior: Scaling Generative Control for Physics-Based Human-Object Interactions
- Kimodo: Scaling Controllable Human Motion Generation
- OmniH2O: Universal and Dexterous Human-to-Humanoid Whole-Body Teleoperation and Learning
- TopoRetarget: Interaction-Preserving Retargeting for Dexterous Manipulation
- A Minimalist Retargeting-Guided Reinforcement Learning Recipe for Dexterous Manipulation
- SynManDex: Synthesizing Human-like Dexterous Grasps from Synthetic Human Pre-Grasps
- Learning Dexterous Manipulation Using Contact Wrench Guidance From Human Demonstration
- ARDY: Autoregressive Diffusion with Hybrid Representation for Interactive Human Motion Generation
- Hand-Eye Autonomous Delivery: Learning Humanoid Navigation, Locomotion and Reaching
- Locate 3D: Real-World Object Localization via Self-Supervised Learning in 3D
- COMPASS: Cross-embodiment Mobility Policy via Residual RL and Skill Synthesis
- VIRAL: Visual Sim-to-Real at Scale for Humanoid Loco-Manipulation
- VAIC: Vision-Guided Humanoid Agile Object Interaction Control via Decoupled Commands
- PhysHSI: Towards a Real-World Generalizable and Natural Humanoid-Scene Interaction System
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving