rMuscle: Robotic Muscle Memory for Efficient Vision-Language-Action Model Inference
cs.RO, cs.AI
Submitted: 2026-09-16
Updated: 2026-09-30
Code: https://github.com/black-forest-labs/flux
License: http://creativecommons.org/licenses/by/4.0/
The gist: Factory work is a promising early scenario for embodied AI: assigning repetitive manual jobs to robots has clear economic payoff, and a structured station keeps the jobs tractable for current
Terminology
Abstract
Factory work is a promising early scenario for embodied AI: assigning repetitive manual jobs to robots has clear economic payoff, and a structured station keeps the jobs tractable for current policies. Vision-Language-Action (VLA) models now dominate as the policy paradigm for these robots. The inference latency of VLA models directly affects robot responsiveness and motion smoothness. However, existing VLA inference frameworks do not fully exploit the characteristics of embodied workloads or account for the distinct bottlenecks across different stages of VLA inference. In this paper, we first characterize embodied workloads and identify substantial task similarity across repeated robot executions. We further find that such similarity extends beyond observations and action trajectories to internal model states. Drawing on these observations, we present rMuscle, a real-time VLA inference framework inspired by human muscle memory. It exploits cross-execution similarity through a dual-phase muscle-memory cache. The Context Cache reuses visual-token outputs to reduce computation, while the Action Cache reuses neuron activation patterns to reduce weight accesses. We keep both the cache memory footprint and access overhead low through online cache recomputation, sliding-window cache retrieval, and mask sharing across consecutive denoising steps. rMuscle achieves 1.29-1.42X speedup on RTX 4090 and Jetson Thor across LIBERO, RoboTwin, and physical manipulation tasks, while maintaining the original success rates on real-world robots.
Sources
- Efficient Vision-Language-Action Models for Embodied Manipulation: A Systematic Survey
- D'ej\`a Vu: Efficient Video-Language Query Engine with Learning-based Inter-Frame Computation Reuse
- How Fast Can I Run My VLA? Demystifying VLA Inference Performance with VLA-Perf
- CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation
- TS-DP: Reinforcement Speculative Decoding For Temporal Adaptive Diffusion Policy Acceleration
- LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning
- RoboMamba: Efficient Vision-Language-Action Model for Robotic Reasoning and Manipulation
- Running VLAs at Real-time Speed
- RoboTwin: Dual-Arm Robot Benchmark with Generative Digital Twins (early version)
- vla.cpp: A Unified Inference Runtime for Vision-Language-Action Models
- Realtime-VLA FLASH: Speculative Inference Framework for Diffusion-based VLAs
- GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
- $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
- Real-Time Execution of Action Chunking Flow Policies
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
- GR-3 Technical Report
- StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing
- $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
- SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics
- Execution-State Capsules: Graph-Bound Execution-State Checkpoint and Restore for Low-Latency, Small-Batch, On-Device Physical-AI Serving
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving