EagleVLA: Towards Onboard Real-Time Robot Control via Foresight-Aligned Asynchronous Inference
cs.RO, cs.AI
Submitted: 2026-07-14
Updated: 2026-10-08
Comments: CoRL 2026
Code: https://github.com/PKU-SEC-Lab/Jetson-PI
License: http://creativecommons.org/licenses/by/4.0/
The gist: Vision-Language-Action (VLA) models have achieved impressive performance on diverse embodied tasks.
Terminology
Abstract
Vision-Language-Action (VLA) models have achieved impressive performance on diverse embodied tasks. However, deploying VLA models on low-power onboard devices, such as the Jetson Orin, remains challenging due to their high computational complexity, which leads to substantial inference latency and low control frequency. Asynchronous inference can partially mask this latency by parallelizing action execution and subsequent inference, but it introduces two critical issues: perception-execution misalignment and long reaction time. In this paper, we propose Jetson-PI, a method for efficient VLA deployment on onboard devices via Foresight-Aligned Asynchronous Correction. To address misalignment, we train a lightweight future correction module that predicts future environment representation conditioned on committed actions, enabling the action expert to directly predict actions from the future time step. To reduce reaction time, we introduce confidence-based scheduling optimization that adaptively balances VLM and action expert invocations, complemented by system-level accelerations including CUDA graph reuse, GPU-resident intermediate buffering, and flow unrolling. Extensive experiments demonstrate that Jetson-PI achieves 8.66x and 5.41x improvements in control frequency compared with naive PyTorch and vla.cpp on NVIDIA Jetson Orin, while outperforming VLASH by 14.8% in average success rate on the LIBERO benchmark. The code of our asynchronous algorithm is available on https://github.com/PKU-SEC-Lab/Jetson-PI, and our efficient llama.cpp-based inference engine is available on https://github.com/PKU-SEC-Lab/Jetson-PI-Edge.
Sources
- $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
- OpenVLA: An Open-Source Vision-Language-Action Model
- $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
- Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success
- DySL-VLA: Efficient Vision-Language-Action Model Inference via Dynamic-Static Layer-Skipping for Robot Manipulation
- CEED-VLA: Consistency Vision-Language-Action Model with Early-Exit Decoding
- Execution-State Capsules: Graph-Bound Execution-State Checkpoint and Restore for Low-Latency, Small-Batch, On-Device Physical-AI Serving
- ForeAct: Steering Your VLA with Efficient Visual Foresight Planning
- VLASH: Real-Time VLAs via Future-State-Aware Asynchronous Inference
- Realtime-VLA FLASH: Speculative Inference Framework for Diffusion-based VLAs
- Realtime-VLA V2: Learning to Run VLAs Fast, Smooth, and Accurate
- KERV: Kinematic-Rectified Speculative Decoding for Embodied VLA Models
- A Survey on Efficient Vision-Language-Action Models
- Efficient Vision-Language-Action Models for Embodied Manipulation: A Systematic Survey
- How Fast Can I Run My VLA? Demystifying VLA Inference Performance with VLA-Perf
- AsyncVLA: An Asynchronous VLA for Fast and Robust Navigation on the Edge
- EdgeVLA: Efficient Vision-Language-Action Models
- RAPID: Redundancy-Aware and Compatibility-Optimal Edge-Cloud Partitioned Inference for Diverse VLA Models
- ActionFlow: A Pipelined Action Acceleration for Vision Language Models on Edge
- A1: A Fully Transparent Open-Source, Adaptive and Efficient Truncated Vision-Language-Action Model
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving