vla.cpp: A Unified Inference Runtime for Vision-Language-Action Models
cs.RO, cs.AI, cs.LG, cs.SY, eess.SY
Submitted: 2026-06-06
Updated: 2026-09-21
Comments: 8 pages, 5 figures. Code available at https://github.com/VinRobotics/vla.cpp
Code: https://github.com/VinRobotics/vla.cpp
Project page: https://fai-modelopt-tech.github.io/vla-cpp.github.io
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics
- Efficient Memory Management for Large Language Model Serving with PagedAttention
- How Fast Can I Run My VLA? Demystifying VLA Inference Performance with VLA-Perf
- OpenVLA: An Open-Source Vision-Language-Action Model
- Diffusion Policy: Visuomotor Policy Learning via Action Diffusion
- Evo-1: Lightweight Vision-Language-Action Model with Preserved Semantic Alignment
- BitVLA: 1-bit Vision-Language-Action Models for Robotics Manipulation
- A Survey on Efficient Vision-Language-Action Models
- LiteVLA-Edge: Quantized On-Device Multimodal Control for Embedded Robotics
- Lite VLA: Efficient Vision-Language-Action Control on CPU-Bound Edge Robots
- NanoVLA: Routing Decoupled Vision-Language Understanding for Nano-sized Generalist Robotic Policies
- VLAgents: A Policy Server for Efficient VLA Inference
- OxyGen: Unified KV Cache Management for VLA Inference under Multi-Task Parallelism
- Characterizing VLA Models: Identifying the Action Generation Bottleneck for Edge AI Architectures
- Characterizing Vision-Language-Action Models across XPUs: Constraints and Acceleration for On-Robot Deployment
- ActionFlow: A Pipelined Action Acceleration for Vision Language Models on Edge
- SpecPrune-VLA: Accelerating Vision-Language-Action Models via Action-Aware Self-Speculative Pruning
- Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success
- Evaluating Real-World Robot Manipulation Policies in Simulation
- LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving