DriftOPD: Sequence-Level Reverse-KL Distillation for One-Step VLA Policies
cs.RO, cs.AI, cs.CV, cs.LG
Submitted: 2026-09-29
Updated: 2026-09-29
Project page: https://yj-jun.github.io/DriftOPD/ABSTRACT
Terminology
Sources
- GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
- Mean-Flow based One-Step Vision-Language-Action
- Target Score Matching
- Generative Modeling via Drifting
- Flow-OPD: On-Policy Distillation for Flow Matching Models
- LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models
- Drift Q-Learning
- Score and Distribution Matching Policy: Advanced Accelerated Visuomotor Policies via Matched Distillation
- Soft Truncation: A Universal Training Technique of Score-based Diffusion Model for High Precision Score Estimation
- RLDX-1 Technical Report
- VLA-RFT: Vision-Language-Action Reinforcement Fine-tuning with Verified Rewards in World Simulators
- DiffusionOPD: A Unified Perspective of On-Policy Distillation in Diffusion Models
- One-Step Flow Policy: Self-Distillation for Fast Visuomotor Policies
- VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning
- DriftWorld: Fast World Modeling through Drifting
- Consistency Policy: Accelerated Visuomotor Policies via Consistency Distillation
- SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics
- VGAS: Value-Guided Action-Chunk Selection for Few-Shot Vision-Language-Action Adaptation
- ABot-M0: VLA Foundation Model for Robotic Manipulation with Action Manifold Learning
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving