vla.simd: Efficient CPU Inference for Language-Conditioned Manipulation
cs.RO, cs.AI, cs.SY, eess.SY
Submitted: 2026-09-21
Updated: 2026-09-21
Comments: 8 pages, 7 tables, 5 figures. Project page: https://vla-simd.github.io/
Code: https://github.com/google/XNNPACK
Project page: https://vla-simd.github.io
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics
- $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
- TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM
- NanoVLA: Routing Decoupled Vision-Language Understanding for Nano-sized Generalist Robotic Policies
- BitVLA: 1-bit Vision-Language-Action Models for Robotics Manipulation
- A Survey on Efficient Vision-Language-Action Models
- How Fast Can I Run My VLA? Demystifying VLA Inference Performance with VLA-Perf
- Characterizing VLA Models: Identifying the Action Generation Bottleneck for Edge AI Architectures
- Characterizing Vision-Language-Action Models across XPUs: Constraints and Acceleration for On-Robot Deployment
- Lite VLA: Efficient Vision-Language-Action Control on CPU-Bound Edge Robots
- Decoupling Vision, Language, and Action for Efficient Multi-Task Robot Policies
- ExecuTorch -- A Unified PyTorch Solution to Run AI Models On-Device
- vla.cpp: A Unified Inference Runtime for Vision-Language-Action Models
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving