VGFM: Expressive Robot Policies via Dense Value Guidance in Flow Matching
cs.RO, cs.LG
Submitted: 2026-09-13
Updated: 2026-09-13
Comments: IROS 2026
License: http://creativecommons.org/licenses/by/4.0/
The gist: Recent robot learning paradigms increasingly rely on large offline datasets of robotic interactions to train control policies.
Terminology
Abstract
Recent robot learning paradigms increasingly rely on large offline datasets of robotic interactions to train control policies. Expressive generative models enable rich and multimodal action representations, expanding the capability of this paradigm for complex robotic control. However, policy improvement with multi-step generative actors remains challenging. In offline reinforcement learning (RL), incorporating value-based objectives along generative trajectories often introduces substantial training complexity, including backpropagation through time (BPTT), auxiliary architectures, or distillation losses. We propose Value-Guided Flow Matching (VGFM), a scalable offline RL framework that enables dense value-guided shaping within a flow-based policy while avoiding BPTT and additional algorithmic overhead. VGFM parameterizes the policy as a conditional flow-matching model in action (x-prediction) space, ensuring that each intermediate flow step produces a valid robot action that can be directly evaluated by a standard offline RL critic. This design allows value guidance to be applied at randomly sampled flow times without differentiating through the entire generative trajectory, while preserving inference-time flexibility by varying the discretization of the underlying flow ODE without retraining. Evaluated on robotic locomotion and manipulation tasks in OGBench, VGFM achieves strong performance across a wide range of tasks under rigorous evaluation protocols. With minimal hyperparameter tuning, these results demonstrate that VGFM provides a simple, scalable, and effective approach for expressive policy learning in long-horizon, goal-oriented robotic control.
Sources
- Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems
- Diffusion Policy: Visuomotor Policy Learning via Action Diffusion
- Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow
- Flow Matching for Generative Modeling
- Diffusion Policies as an Expressive Policy Class for Offline Reinforcement Learning
- Diffusion-Based Planning for Autonomous Driving with Flexible Guidance
- Flow Q-Learning
- Scaling Offline RL via Efficient and Expressive Shortcut Models
- Flow-Based Single-Step Completion for Efficient and Expressive Policy Learning
- OGBench: Benchmarking Offline Goal-Conditioned RL
- Behavior Regularized Offline Reinforcement Learning
- Score-Based Generative Modeling through Stochastic Differential Equations
- Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning
- Consistency Models as a Rich and Efficient Policy Class for Reinforcement Learning
- Score Regularized Policy Optimization through Diffusion Behavior
- Actor-Critic without Actor
- Back to Basics: Let Denoising Generative Models Denoise
- Offline Reinforcement Learning with Implicit Q-Learning
- IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving