PhyGile: Physics-Prefix Guided Motion Generation for Agile General Humanoid Motion Tracking
cs.RO, cs.AI, cs.CV
Submitted: 2026-03-13
Updated: 2026-08-30
License: http://creativecommons.org/licenses/by/4.0/
The gist: Humanoid robots are expected to execute agile and expressive whole-body motions in real-world settings.
Terminology
Abstract
Humanoid robots are expected to execute agile and expressive whole-body motions in real-world settings. Existing text-to-motion generation models are predominantly trained on captured human motion datasets, whose priors assume human biomechanics, actuation, mass distribution, and contact strategies. When such motions are directly retargeted to humanoid robots, the resulting trajectories may satisfy geometric constraints (e.g., joint limits and pose continuity) and appear kinematically reasonable. However, they frequently violate the physical feasibility required for real-world execution. To address these issues, we present PhyGile, a unified framework that closes the loop between robot-native motion generation and General Motion Tracking (GMT). PhyGile performs physics-prefix-guided robot-native motion generation at inference time, directly generating robot-native motions in a 262-dimensional skeletal space with physics-guided prefixes, thereby eliminating inference-time retargeting artifacts and reducing generation-execution discrepancies. Before physics-prefix adaptation, we train the GMT controller with a curriculum-based mixture-of-experts scheme, followed by post-training on unlabeled motion data to improve robustness over large-scale robot motions. During physics-prefix adaptation, the GMT controller is further fine-tuned with generated objectives under physics-derived prefixes, enabling agile and stable execution of complex motions on real robots. Extensive offline and real-robot experiments demonstrate that PhyGile expands the frontier of text-driven humanoid control, enabling stable tracking of agile, highly difficult whole-body motions that go well beyond walking and low-dynamic motions typically achieved by prior methods.
Sources
- GMT: General Motion Tracking for Humanoid Whole-Body Control
- EGM: Efficiently Learning General Motion Tracking Policy for High Dynamic Humanoid Whole-Body Control
- General Humanoid Whole-Body Control via Pretraining and Fast Adaptation
- Human Motion Diffusion Model
- TextOp: Real-time Interactive Text-Driven Humanoid Robot Motion Generation and Control
- UniAct: Unified Motion Generation and Action Streaming for Humanoid Robots
- LangWBC: Language-directed Humanoid Whole-Body Control via End-to-end Learning
- Retargeting Matters: General Motion Retargeting for Humanoid Motion Tracking
- SENTINEL: A Fully End-to-End Language-Action Model for Humanoid Whole Body Control
- Human Motion Diffusion as a Generative Prior
- CLoSD: Closing the Loop between Simulation and Diffusion for multi-task character control
- AvatarCLIP: Zero-Shot Text-Driven Generation and Animation of 3D Avatars
- Humanoid-VLA: Towards Universal Humanoid Control with Visual Integration
- Commanding Humanoid by Free-form Language: A Large Language Action Model with Unified Motion Vocabulary
- From Language to Locomotion: Retargeting-free Humanoid Control via Motion Latent Guidance
- Do You Have Freestyle? Expressive Humanoid Locomotion via Audio Control
- SONIC: Supersizing Motion Tracking for Natural Humanoid Whole-Body Control
- Hume: Introducing System-2 Thinking in Visual-Language-Action Model
- MOSAIC: Bridging the Sim-to-Real Gap in Generalist Humanoid Motion Tracking and Teleoperation with Rapid Residual Adaptation
- From Experts to a Generalist: Toward General Whole-Body Control for Humanoid Robots
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving