SafeFlow: Real-Time Text-Driven Humanoid Whole-Body Control via Physics-Guided Rectified Flow and Selective Safety Gating
cs.RO, cs.AI, cs.SY, eess.SY
Submitted: 2026-03-25
Updated: 2026-09-15
Comments: Project Page: https://hanbyelcho.info/safeflow/
Project page: https://hanbyelcho.info/safeflow
License: http://creativecommons.org/licenses/by-sa/4.0/
The gist: Recent advances in real-time interactive text-driven motion generation have enabled humanoids to perform diverse behaviors.
Terminology
Abstract
Recent advances in real-time interactive text-driven motion generation have enabled humanoids to perform diverse behaviors. However, kinematics-only generators often exhibit physical hallucinations, producing motion trajectories that are physically infeasible to track with a downstream motion tracking controller or unsafe for real-world deployment. These failures often arise from the lack of explicit physics-aware objectives for real-robot execution and become more severe under out-of-distribution (OOD) user inputs. Hence, we propose SafeFlow, a text-driven humanoid whole-body control framework that combines physics-guided motion generation with a 3-Stage Safety Gate driven by explicit risk indicators. SafeFlow adopts a two-level architecture. At the high level, we generate motion trajectories using Physics-Guided Rectified Flow Matching in a VAE latent space to improve real-robot executability, and further accelerate sampling via Reflow to reduce the number of function evaluations (NFE) for real-time control. The 3-Stage Safety Gate enables selective execution by detecting semantic OOD prompts using a Mahalanobis score in text-embedding space, filtering unstable generations via a directional sensitivity discrepancy metric, and enforcing final hard kinematic constraints such as joint and velocity limits before passing the generated trajectory to a low-level motion tracking controller. Extensive experiments on the Unitree G1 demonstrate that SafeFlow outperforms diffusion- and retargeting-based baselines in success rate, physical compliance, and inference speed while preserving motion diversity, with consistent gains across three downstream tracking controllers.
Sources
- MotionDiffuse: Text-Driven Human Motion Generation with Diffusion Model
- TextOp: Real-time Interactive Text-Driven Humanoid Robot Motion Generation and Control
- Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow
- GaussGym: An open-source real-to-sim framework for learning locomotion from pixels
- A Unified and General Humanoid Whole-Body Controller for Versatile Locomotion
- Humanoid Whole-Body Locomotion on Narrow Terrain via Dynamic Balance and Reinforcement Learning
- Opening the Sim-to-Real Door for Humanoid Pixel-to-Action Policy Transfer
- VIRAL: Visual Sim-to-Real at Scale for Humanoid Loco-Manipulation
- Visual Imitation Enables Contextual Humanoid Control
- GMT: General Motion Tracking for Humanoid Whole-Body Control
- BeyondMimic: From Motion Tracking to Versatile Humanoid Control via Guided Diffusion
- Track Any Motions under Any Disturbances
- SONIC: Supersizing Motion Tracking for Natural Humanoid Whole-Body Control
- KungfuBot2: Learning Versatile Motion Skills for Humanoid Whole-Body Control
- TWIST: Teleoperated Whole-Body Imitation System
- TWIST2: Scalable, Portable, and Holistic Humanoid Data Collection System
- PhysHOI: Physics-Based Imitation of Dynamic Human-Object Interaction
- LangWBC: Language-directed Humanoid Whole-Body Control via End-to-end Learning
- Learning Transferable Visual Models From Natural Language Supervision
- Isaac Lab: A GPU-Accelerated Simulation Framework for Multi-Modal Robot Learning
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving