Action with Visual Primitives
cs.RO, cs.AI
Submitted: 2026-05-21
Updated: 2026-09-20
Comments: We wish to complete further experimental improvements before publishing this work
Project page: https://kingdroper.github.io/AVP
License: http://creativecommons.org/licenses/by/4.0/
The gist: Vision-Language-Action (VLA) models have emerged as a promising paradigm for generalist robotic manipulation.
Terminology
Abstract
Vision-Language-Action (VLA) models have emerged as a promising paradigm for generalist robotic manipulation. A common design in current architectures maps language instructions and visual observations to actions in a single forward pass. While conceptually simple, this formulation entangles instruction comprehension, spatial scene understanding, and motor control within a single learning objective. As a result, the action expert must implicitly relearn cognitive and perceptual capabilities already present in the pretrained VLM, which can limit both learning efficiency and generalization. We introduce AVP (Action with Visual Primitives), an end-to-end architecture that implements this visual-primitive-centric interface: the VLM infers the next-stage target and emits visual-primitive tokens that condition a flow-matching action expert, with supervision derived from end-effector kinematics. Real-robot experiments on general pick-and-place tasks show that AVP improves the success rate by 37.04% over pi 0.5 and outperforms other recent methods, with consistent gains in data efficiency, spatial-compositional generalization, and object-level transfer.
Sources
- OpenVLA: An Open-Source Vision-Language-Action Model
- $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
- $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
- RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation
- Octo: An Open-Source Generalist Robot Policy
- Open X-Embodiment: Robotic Learning Datasets and RT-X Models
- DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset
- RoboVerse: Towards a Unified Platform, Dataset and Benchmark for Scalable and Generalizable Robot Learning
- VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action Models
- Don't Blind Your VLA: Aligning Visual Representations for OOD Generalization
- Actions as Language: Fine-Tuning VLMs into VLAs Without Catastrophic Forgetting
- RT-H: Action Hierarchies Using Language
- Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models
- ${\pi}_{0.7}$: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities
- Vista: A Generalizable Driving World Model with High Fidelity and Versatile Controllability
- Fast-WAM: Do World Action Models Need Test-time Future Imagination?
- Point What You Mean: Visually Grounded Instruction Policy
- VP-VLA: Visual Prompting as an Interface for Vision-Language-Action Models
- TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies
- GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving