SPIDER: Scalable Physics-Informed Dexterous Retargeting
cs.RO, cs.CV
Submitted: 2025-11-12
Updated: 2026-09-26
Comments: Project website: https://jc-bao.github.io/spider-project/
Code: https://github.com/thu-ml/RDT2
Project page: https://jc-bao.github.io/spider-project
License: http://creativecommons.org/licenses/by/4.0/
The gist: Learning agile robotic policies requires large-scale demonstrations, but translating abundant human motion data to robots is bottlenecked by the embodiment gap and missing dynamic information.
Terminology
Abstract
Learning agile robotic policies requires large-scale demonstrations, but translating abundant human motion data to robots is bottlenecked by the embodiment gap and missing dynamic information. To bridge this gap, we propose Scalable Physics-Informed DExterous Retargeting (SPIDER), a physics-based retargeting framework to transform and augment kinematic-only human demonstrations into dynamically feasible robot trajectories at scale. Our key insight is that human demonstrations should provide global task structure and objective, while a sampling-based solver can be used to find the feasible solution given the physics constraints. As a general framework, SPIDER is an efficient physics-based retargeting method that can be applied to both humanoid whole-body loco-manipulation and dexterous manipulation across 9 humanoid/dexterous hand embodiments, 6 datasets and 3 simulators. By bypassing the need for policy optimization, it achieves state-of-the-art performance while being 10x faster than reinforcement learning (RL) baselines. Furthermore, SPIDER enables physics-based data augmentation, such as imposing external payloads, perturbations, and new contact patterns, to generate diverse data. We demonstrate that the retargeted motion can be executed on the robot directly open-loop or serve as feasible reference motion for efficient RL policy learning. Our pipeline enables large-scale generation of robot trajectories from human motion, and we release a dataset of 6,520 simulated demonstrations across four hands and 191 dataset-specific object identities to support future research.
Sources
- Retargeting Matters: General Motion Retargeting for Humanoid Motion Tracking
- A Tutorial on Bayesian Optimization
- EgoDex: Learning Dexterous Manipulation from Large-Scale Egocentric Video
- Predictive Sampling: Real-time Behaviour Synthesis with MuJoCo
- BeyondMimic: From Motion Tracking to Versatile Humanoid Control via Guided Diffusion
- EgoZero: Robot Learning from Smart Glasses
- DexMachina: Functional Retargeting for Bimanual Dexterous Manipulation
- MimicGen: A Data Generation System for Scalable Robot Learning using Human Demonstrations
- DemoDiffusion: One-Shot Human Imitation using pre-trained Diffusion Policy
- Physics-based Motion Retargeting from Sparse Inputs
- Evolution Strategies as a Scalable Alternative to Reinforcement Learning
- ExoStart: Efficient learning for dexterous manipulation with sensorized exoskeleton demonstrations
- HDMI: Learning Interactive Humanoid Whole-Body Control from Human Videos
- One-Shot Transfer of Long-Horizon Extrinsic Manipulation Through Contact Retargeting
- EO-1: An Open Unified Embodied Foundation Model for General Robot Control
- Proximal Policy Optimization Algorithms
- OmniRetarget: Interaction-Preserving Data Generation for Humanoid Whole-Body Loco-Manipulation and Scene Interaction
- DexGraspVLA: A Vision-Language-Action Framework Towards General Dexterous Grasping
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving