Premover: Fast Vision-Language-Action Control via Early Execution During Instruction Delivery
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: I'm Rosa, and with me are Dev and Taro, guest researcher.
Dev: Today's paper: "Premover: Fast Vision-Language-Action Control via Early Execution During Instruction Delivery".
Rosa: Vision-Language-Action (VLA) policies are typically evaluated as if the user has finished typing or speaking before the robot begins acting, but this assumption ignores significant idle time during instruction input.
Dev: First, who's behind it and why it matters.
Paper summary: Rosa: So, we're talking about the paper "Premover: Fast Vision-Language-Action Control via Early Execution During Instruction Delivery," and the main idea is that current Vision-Language-Action policies are being evaluated incorrectly because they assume the user is done speaking before the robot even starts acting.
Dev: That idle time during instruction input is a big issue, and Premover proposes a way to use that waiting period for useful precomputation by allowing the VLA backbone to start acting earlier based on what's already been typed or spoken.
Taro: I'm interested in how this affects the system when things don't go according to plan, Rosa; if we act early based on partial instructions, what happens when the world throws a curveball that contradicts our current understanding of the task?
Rosa: That’s a fair point, Taro; it brings up the challenge of where to focus visually when you only have a partial instruction, as Premover has to handle that without letting the backbone look everywhere.
Dev: Exactly, and how does it solve that focusing problem while also deciding when to commit to an action?
Taro: I'm curious about the readiness mechanism; if the system is acting based on a partial prefix, what tells it when that prefix is specific enough to warrant committing to a real target?
Rosa: Well, Premover uses a focus map derived from comparing image patches against language tokens in a shared space, which then gets reweighted at each step.
Dev: That focus map acts as an input reweighting mechanism controlled by a floor scale parameter via those weight functions, which is pretty clever for managing the flow of information.
Taro: So, the readiness gate essentially measures how concentrated that focus map has become and compares it to a threshold to decide whether to execute an action at time t?
Rosa: Precisely; Premover executes an action at time t if the readiness score rt exceeds that learned threshold tau, which means it waits until the prefix has localized a sufficiently specific visual target.
Dev: And I see how that directly addresses the latency issue because it allows forward passes during input, and the paper shows this reduces end-to-end latency to eighty-six point four percent of the full-prompt baseline on LIBERO <ref:2605.12160#pg1>.
Taro: That reduction in time sounds significant for real deployment scenarios where you can't wait for a user to finish everything before starting movement, but does this early execution still run into issues with failure modes, like executing an action based on a misunderstanding of the partial instruction?
Rosa: The paper addresses those concerns by supervising the focus map using simulator-rendered target-object segmentation masks and also having a readiness supervision term that signals when it's too early to act.
Paper summary: Dev: That dual supervision approach with Lfocus and Lready, limited to only the two projection heads fimg and flang, keeps the trainable parameters small—less than one percent of the backbone's parameter count—which is good for efficiency <ref:2605.12160#pg0>.
Taro: So, if we look at the overall impact of this Premover module, what are we really looking at in terms of how this technology might shape future autonomous systems outside of a controlled lab setting?
Rosa: It suggests that VLA policies don't have to be waiting around for perfect input; they can be proactive in using partial information to make decisions faster.
Dev: The results on the LIBERO benchmark showed a reduction in mean wall-clock time from thirty-four point zero seconds down to twenty-nine point four seconds, which is about a thirteen point six percent reduction, matching the success rate of the full-prompt baseline at ninety-five point one percent versus ninety-five point zero percent <ref:2605.12160#pg1>.
Taro: That comparison between Premover's performance and naive premoving collapsing to a sixty-six point four percent success rate shows that this mechanism actually helps maintain performance when things get tricky, which is important for real-world robustness <ref:2605.12160#pg1>.
Rosa: It really shows that capturing input latency without the success collapse of unconstrained streaming is what makes this approach compelling for field robotics. So, how does this early execution strategy change how we think about instruction delivery altogether?
Taro: Thinking about the broader implications, if Premover can effectively handle partial instructions and decide when to act based on localized focus, it opens up possibilities for much more responsive and interactive physical systems.
Dev: From an engineering standpoint, the fact that they've managed to keep the trainable parameters extremely limited while achieving this latency reduction is a testament to how constrained we can keep these specialized modules.
Rosa: I wonder if this means we could see these types of systems operating in environments with very noisy or incomplete instructions more effectively than before?
Taro: That's where the robustness comes in; if the readiness gate correctly filters out premature actions based on insufficient visual information, it suggests a system that can handle ambiguity better during instruction streams <ref:2605.12160#pg1>.
Dev: It definitely moves the decision-making process closer to real-time interaction rather than waiting for a complete command before processing begins.
Rosa: So, this paper, "Premover: Fast Vision-Language-Action Control via Early Execution During Instruction Delivery," essentially gives us a framework to make VLA policies operate more intelligently during the instruction input phase by using partial data to guide early action initiation.
Taro: And I think the core contribution lies in coupling the focus map mechanism for where to look with the readiness gate for when to act, which tackles those two coupled challenges mentioned in the motivation section.
Paper summary: Dev: That coupling is key; it's not just about looking at things, it's about having a learned strategy for committing to action based on the stream itself.
Rosa: It moves us away from treating the instruction input time as dead time and towards using that window as active precomputation time for the robot.
Taro: If this works well in simulation, I think we could see these systems deployed in more complex physical tasks where instructions might be delivered incrementally or conversationally.
Dev: The paper's evaluation on VLA-arena showed a ten point three percent reduction in wall-clock time with a 2 point 1 percentp success rate gap compared to naive premoving, confirming that the readiness gate captures input-time latency without the success collapse of unconstrained streaming <ref:2605.12160#pg1>.
Rosa: It really validates that this mechanism works for real operational scenarios, and I'm excited to see how this translates to actual field robotics where we can't always guarantee a perfect instruction upfront.
Dev: The limitation the authors point out is that Premover targets a complementary regime, meaning it's designed for grounding partial instructions as they stream in, which implies its primary strength lies in this specific interaction style rather than handling completely novel instruction formats from scratch <ref:2605.12160#pg2>.
Taro: So while it's strong at streaming prefixes, we still need to figure out how it generalizes when the instruction structure is entirely different or very sparse.
Rosa: That points toward future work needing to explore how this module adapts beyond just the specific regime of streaming prefixes that Premover targets.
Dev: For now, the immediate impact is a more responsive control loop that minimizes idle time, which directly translates to lower latency in execution and better real-time interaction capability for physical robots <ref:2605.12160#pg1>.
Taro: It’s about making the system feel less sluggish during instruction reception, which is a practical improvement for any autonomous agent interacting with the physical world.
Rosa: So, to wrap up this discussion on "Premover: Fast Vision-Language-Action Control via Early Execution During Instruction Delivery," we've seen how it uses a focus map and readiness gate to convert idle time into useful precomputation by acting on partial instructions <ref:2605.12160#pg0>.
Dev: That system is designed to reduce end-to-end latency significantly, achieving results like the eighty-six point four percent of the full-prompt baseline on LIBERO and showing improvements in other benchmarks <ref:2605.12160#pg1>.
Taro: The implication for autonomy is that we can build systems that are more proactive during instruction delivery, handling ambiguity by deciding when to commit based on the visual evidence gathered so far <ref:2605.12160#pg1>.
Rosa: It seems like Premover gives us a concrete way to manage the inherent time gap between receiving an instruction and starting physical action in a way that respects the streaming nature of human input.
Conclusion: Rosa: So, we've seen how Premover uses a focus map and a readiness gate to convert idle time into useful precomputation by acting on partial instructions during delivery. Dev, what are your thoughts on the title and the authors of this paper?
Dev: I think "Premover" is pretty descriptive; it really captures that idea of using early execution for faster control loops. The authors clearly focused on minimizing latency, which is crucial for any real-time system we build.
Taro: I agree with Dev on the latency focus; it's all about getting that loop rate up without having to wait for the whole input sequence to finish before starting motion.
Rosa: Exactly, and looking at the authors, they seem like people who understand how to get these complex VLA systems running more efficiently in practice. What does this mean for us out there in the field?
Dev: It means we can deploy robots that feel much more responsive during a conversation or instruction stream rather than having them freeze up waiting for the whole command. We're talking about reducing that thirty-nine percent idle time they mentioned, which is a big deal for operational efficiency.
Taro: And I see it as a way to make autonomy more interactive; if the robot can start moving based on what it understands so far, it feels like a much more natural interaction with humans or the environment.
Rosa: It really suggests that we're moving away from waiting for perfect input and toward using whatever we have right now to make progress. So, how does this concept of early execution fit into the broader picture of autonomous agents interacting with people?
UNIST · The Catholic University of Korea
cs.RO, cs.AI
Submitted: 2026-05-12
Updated: 2026-10-05
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 88/100
The gist: Vision-Language-Action (VLA) policies are typically evaluated as if the user has finished typing or speaking before the robot begins acting, but this assumption ignores significant idle time during
Key concepts
- Streaming Prefix
- This refers to the partial instruction (text or speech) received so far during an interaction. Premover allows the system to begin acting based on this incomplete information instead of waiting for the entire instruction to be complete, effectively starting computation earlier.
- Focus Map
- A map generated by comparing every image patch against the tokens in the streaming prefix within a shared latent space. It shows which visual areas are relevant based on what has been heard or seen previously, guiding where the policy should look next.
- Readiness Gate
- A learnable mechanism that decides whether to start executing an action. It measures how concentrated the focus map is and compares this concentration against a learned threshold. The robot only acts if this readiness score exceeds the threshold, ensuring it waits for a sufficiently specific visual target.
- Projection Heads
- Two small modules (one for images, one for language) that map an intermediate layer of the frozen VLA backbone into a shared latent space. They create the necessary coordinates to compare visual and linguistic information effectively to generate the focus map.
Terminology
Summary
Vision-Language-Action (VLA) policies are typically evaluated as if the user has finished typing or speaking before the robot begins acting, but this assumption ignores significant idle time during instruction input. Premover introduces a lightweight module that converts this idle window into useful precomputation by allowing the VLA backbone to begin acting earlier based on partial instructions.
The gist
Premover is a lightweight module that converts the idle window into useful precomputation by keeping the VLA backbone frozen and attaching two small projection heads—one for image patches, one for language tokens—that map an intermediate layer of the backbone into a shared space to create a focus map, which is then used with a learnable readiness gate to decide when the policy should begin acting.
Motivation and Challenges
Conventional VLA inference assumes that the full instruction (i.e., query) is available before the policy begins computation, leaving it idle for a substantial fraction of interaction time—averaging 39% across LIBERO suites. The paper argues that this interval should not be treated as dead time and proposes acting on the partial instruction observed so far, referred to as the streaming prefix.
This approach faces two coupled challenges: (i) where to attend under partial language, as a frozen VLA backbone's vision-language alignment may spread attention across irrelevant regions; and (ii) when to start acting, as acting before the prefix has identified the referent risks committing the robot to a wrong target.
How it works
Premover consists of two complementary components: a focus map and a readiness gate, both computed from a shared latent space using two small modality-specific projection heads—one for image patches and one for language tokens—that map an intermediate layer of the frozen backbone into the same coordinates. The focus map is formed by comparing each image patch against the streaming prefix tokens in this shared space, yielding a per-patch probability, denoted as the focus map. This map is then used as an input reweighting mechanism at the next step: we instead reuse the focus map from step t as the input reweighting at step t+1.
This reweighting is controlled by a floor scale parameter, parameterized by a weight function:
(4) and (5)
The readiness gate decides when to start acting. It measures how concentrated the focus map has become and compares this concentration to a single learnable threshold, denoted as τ. The readiness score is computed as:
(6)
The policy executes an action at time t if the readiness score rt exceeds the threshold τ: execute action at t ⇐⇒ rt ≥ τ.
This mechanism ensures that the policy begins acting only once the prefix has localized a sufficiently specific visual target.
Training Objective and Supervision
The two components are jointly trained via a weighted sum of two losses:
(9)
-
Focus Map Supervision: A class-balanced per-patch Binary Cross-Entropy (BCE) loss, Lfocus, is applied to supervise the focus map with simulator-rendered target-object segmentation masks (Eq. 3). The loss is supervised only when the target has already emerged in the prefix.
-
Readiness Supervision: A BCE loss, Lready, supervises the readiness gate using a logit formed by dividing the gap between the readiness score and threshold by a temperature T > 0 (Eq. 8). This term is used for prefixes where
the target has not yet emerged,
signaling whenit is too early.
The trainable parameters are limited to the two projection heads (fimg and flang) and the readiness threshold τ, which account for less than 1% of the backbone's parameter count.
Evaluation and Results
Premover was evaluated on π0.5 on LIBERO and VLA-arena benchmarks. On LIBERO, Premover reduced mean wall-clock time from 34.0 to 29.4 seconds, a 13.6% reduction,
while matching the full-prompt baseline’s success rate (95.1% vs. 95.0%). Naive premoving collapsed to a success rate of 66.4%. On VLA-arena, Premover reduced wall-clock time by 10.3% with a 2.1%p success rate gap compared to naive premoving, confirming that the readiness gate captures input-time latency without the success collapse of unconstrained streaming.
The ablation study showed that the readiness gate alone recovers 21.9%p of mean success, while adding focus map injection adds another 6.7%p to reach 95.1%.
Improvements for AI systems
Based on the provided scientific paper Premover: Fast Vision-Language-Action Control,
here are specific improvements for existing Vision-Language-Action (VLA) systems and what those improved systems can achieve:
Improvement 1: Implement a Lightweight, Frozen Backbone Modulation Module (Premover)
The core improvement is the introduction of Premover, a lightweight module attached to an existing frozen VLA backbone. This module avoids costly full model retraining by only training two small projection heads and one scalar threshold.
Specific enhancements include:
-
A shared latent space between image patches and language tokens to compute a per-patch focus map.
-
A learnable readiness gate that decides when the policy should commit to an action based on the concentration of this focus map, rather than waiting for the full instruction.
-
Per-patch reweighting of subsequent image tokens using a floor scale parameter (α) to amplify tokens relevant to the streaming prefix while preserving background context for non-target behaviors (like obstacle avoidance).
What improved AI systems can do:
The system will drastically reduce end-to-end wall-clock time during real user interaction. Specifically, it can execute actions significantly earlier—up to 13.6% faster on LIBERO benchmarks—by converting the idle user input
interval into productive precomputation time, leading to a more responsive and interactive human-robot experience.
Improvement 2: Transition from Implicit Attention to Explicit Grounding via Supervised Focus Maps
The system improves visual grounding by explicitly training a focus map (a per-patch relevance distribution) supervised by simulator-rendered target-object segmentation masks.
Specific enhancements include:
-
Using two modality-specific projection heads to map backbone hidden states into a shared space, allowing for direct comparison between visual features and language tokens.
-
A Binary Cross-Entropy loss applied to this focus map during training, ensuring the learned focus map aligns with the true target object location in the simulator.
What improved AI systems can do:
The system will significantly improve robustness against attention dispersion.
Instead of letting a frozen backbone vaguely attend to irrelevant background noise when only a partial instruction is present (as seen in naive premoving), it will concentrate its visual attention precisely on the object indicated by the streaming prefix, leading to higher success rates (up to 98.6% on LIBERO-Object) and better disambiguation of distractors.
Improvement 3: Introduce a Learnable Readiness Criterion for Action Commitment
The system introduces a readiness gate based on an action readiness score (rt), which measures the concentration of the focus map, and compares it to a learnable threshold (τ).
Specific enhancements include:
-
Calculating rt by computing the top-K mean of patch probabilities from the focus map.
-
Training a single scalar threshold τ to gate action execution: act if rt ≥ τ, otherwise hold.
What improved AI systems can do:
The system will solve the critical trade-off between speed and safety/correctness in streaming scenarios. It ensures that the robot commits to an action only when it is confident (i.e., the prefix has localized a referent). This prevents premature commitment
to wrong targets, which naive streaming risks, while still capitalizing on input latency reduction.
Improvement 4: Implement Dynamic Input Scheduling Based on Latency
The system uses a fixed schedule synchronized with the measured per-step inference latency (calibrated to 52.24 WPM) to reveal tokens in the streaming prefix.
Specific enhancements include:
-
Synchronizing token revelation with policy execution steps, ensuring consistent streaming dynamics across all comparisons.
-
Using this synchronization to define an identical input interval for speedup metrics, isolating the effect of Premover from variations in typing speed distribution (as shown by Table 4).
What improved AI systems can do:
The system will provide a standardized and fair evaluation protocol for latency reduction claims across different VLA models, ensuring that performance gains are attributable to the precomputation module rather than variations in how users type.
Abstract
Vision-Language-Action (VLA) policies are typically evaluated under the assumption that the robot starts acting only after the user has finished typing or speaking. In real interactions, however, entering an instruction can take several seconds, leaving the policy idle for a substantial fraction of the interaction. Partial instructions may already contain sufficient information to begin acting before the full instruction arrives. We introduce Premover, a parameter-lightweight module that reduces interaction latency by overlapping instruction delivery with robot execution while keeping the VLA backbone frozen. Premover uses a learned vision-language focus map to identify where the current partial instruction refers in the visual scene, and an action readiness gate to determine when execution can begin. On the SO-101 robot with speech and online transcription, Premover reduces mean wall-clock time by 17.9%, from 60.5s to 49.7s, while maintaining task success rates. In simulation on LIBERO with pi0.5 at the speaking rate, Premover reduces mean wall-clock time by 8.3% while maintaining comparable success to waiting for the complete instruction before execution.
Sources
- $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
- $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
- Bring My Cup! Personalizing Vision-Language-Action Models with Visual Attentive Prompting
- Bridging the Semantic-Action Gap in Visual Token Pruning for Efficient VLA Inference
- GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
- SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics
- ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver
- VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models
- PVI: Plug-in Visual Injection for Vision-Language-Action Models
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving