Veo-Act: Enhancing VLA Policies with Frontier Video Models
cs.RO
Submitted: 2026-04-06
Updated: 2026-09-16
Terminology
Sources
- OpenVLA: An Open-Source Vision-Language-Action Model
- $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
- $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
- Improving Vision-Language-Action Model with Online Reinforcement Learning
- UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent
- HiRT: Enhancing Robotic Control with Hierarchical Robot Transformers
- UniJEPA: Enhancing Robot Policy via Unified Continuous and Discrete Representation Learning
- RT-1: Robotics Transformer for Real-World Control at Scale
- Knowledge Insulating Vision-Language-Action Models: Train Fast, Run Fast, Generalize Better
- Actions as Language: Fine-Tuning VLMs into VLAs Without Catastrophic Forgetting
- VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action Models
- Imagen Video: High Definition Video Generation with Diffusion Models
- Make-A-Video: Text-to-Video Generation without Text-Video Data
- VideoPoet: A Large Language Model for Zero-Shot Video Generation
- Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations
- Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning
- World Action Models are Zero-shot Policies
- VideoVLA: Video Generators Can Be Generalizable Robot Manipulators
- Ctrl-World: A Controllable Generative World Model for Robot Manipulation
- Video Generators are Robot Policies
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving