HybridFlow: A 2-NFE Generative Policy for Real-Time Robotic Manipulation
cs.RO, cs.AI
Submitted: 2026-02-14
Updated: 2026-09-20
Comments: 9 pages. Updated author list and title; expanded analysis, controlled ablations, real-robot evaluation, and VLA experiments. Project page: https://hybridflow-anonymous.pages.dev/
Code: https://github.com/thu-ml/RDT2
Project page: https://hybridflow-anonymous.pages.dev
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Generative policies for robotic manipulation must balance action accuracy with inference latency.
Terminology
Abstract
Generative policies for robotic manipulation must balance action accuracy with inference latency. We present HybridFlow, a three-stage policy inference procedure requiring two network function evaluations (2-NFE). A Global Jump uses the MeanFlow average velocity to generate a coarse action trajectory; a parameter-free ReNoise interpolation constructs a state at a nonzero refinement time; and a Local Refine queries the instantaneous-velocity limit of the same network at that time. This construction reuses a unified model without distillation. Our analysis characterizes interval-composition errors and the attenuation of endpoint error under ReNoise interpolation. Controlled RoboMimic ablations support the MeanFlow proposal and intermediate-state construction, achieving 95% average success with reused noise and 95.5% with fresh noise, versus 78% for one-step MeanFlow. Across five real-robot settings with all policies running on the same Jetson AGX Thor, HybridFlow improves normalized task performance over 16-step Diffusion Policy by 13-68 points with approximately eightfold lower action-generation latency. Additional experiments demonstrate its compatibility as an action expert in a vision-language-action framework. Project page: https://hybridflow-anonymous.pages.dev/
Sources
- 3D Diffusion Policy: Generalizable Visuomotor Policy Learning via Simple 3D Representations
- $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
- $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
- Improved Mean Flows: On the Challenges of Fastforward Generative Models
- Joint Distillation for Fast Likelihood Evaluation and Sampling in Flow-based Models
- Improving the Training of Rectified Flows
- Two-Step Diffusion: Fast Sampling and Reliable Prediction for 3D Keller--Segel and KPP Equations in Fluid Flows
- MeanFlow-TSE: One-Step Generative Target Speaker Extraction with Mean Flow
- RMFlow: Refined Mean Flow by a Noise-Injection Step for Multimodal Generation
- Consistency Models
- Progressive Distillation for Fast Sampling of Diffusion Models
- Improved Techniques for Training Consistency Models
- Consistency Policy: Accelerated Visuomotor Policies via Consistency Distillation
- ManiCM: Real-time 3D Diffusion Policy via Consistency Model for Robotic Manipulation
- One-Step Diffusion Policy: Fast Visuomotor Policies via Diffusion Distillation
- A Magnetic Analog of Pressure-Strain Interaction
- OGPO: One-Step Generative Policy Optimization for Real-Time Robot Control
- Global Well-posedness and Convergence Analysis of Score-based Generative Models via Sharp Lipschitz Estimates
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving