ExecVLA: Following Fine-Grained Execution Constraints in Vision-Language-Action Models with Bi-Level Action Representation
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: Today's paper: "ExecVLA: Following Fine-Grained Execution Constraints in Vision-Language-Action Models with Bi-Level Action Representation".
Dev: The gist The KineVLA framework introduces a kinematics-rich vision-language-action (VLA) task that explicitly decouples goal-level invariance from kinematics-level variability through a bi-level action representation and bi-level reasoning tokens to serve…
Rosa: First, who's behind it and why it matters.
Title and authors: Rosa: So we’re looking at ExecVLA today. This paper is all about taking vision language action models and making them capable of following very specific motion instructions, not just general goals. We're talking about adding a layer that handles the fine details of movement while keeping the overall task objective steady.
Dev: Exactly. The core idea they introduce is this bi-level structure that separates what you want to achieve from exactly how you need to move to get there. It’s about decoupling goal-level invariance from kinematics-level variability, which is a big deal for real robots out in the world.
Taro: I'm curious where this actually gets tested outside of a controlled lab setting. Can we expect this fine-grained control to hold up when the robot encounters unexpected obstacles or real-world noise?
Rosa: That’s what I’ll ask you about later, Taro, but right now, the paper is focusing on how they build this framework. They propose using a bi-level action representation and bi-level reasoning tokens as explicit intermediate variables to connect the language instructions directly to the robot's control signals.
Dev: The mechanism they use is a Bi-Level Residual VQ-VQE system, which means they break down actions into two distinct latent spaces: one for the general goal and one for the specific motion parameters like distance or velocity. That’s how you get those precise motion representations they mentioned.
Taro: So, if the model is generating these two levels of tokens—one coarse and one fine—how does that actually help when things go wrong in execution? What happens when the world misbehaves?
Rosa: The paper suggests using bi-level reasoning tokens to align the internal representations with language parsing at both levels. One level handles the general task goal, and the other specifies those crucial kinematics parameters like direction or orientation, which are annotated in their dataset.
Dev: And they put some mutual information regularization in there to make sure the reasoning text actually matches what’s happening in the action execution. They're maximizing this conditional mutual information, I(Reasoning; Action C), to ensure consistency between what the model is thinking and what it's doing.
Taro: That makes sense for grounding things, but if we look at the results, they show state-of-the-art performance on kinematics-aware benchmarks. What’s the trade-off there? Are we getting better precision for a noticeable hit in speed or complexity?
Rosa: They say that while all methods can achieve similar success rates for reaching the final task goal, when you look at the success rate specifically for following the precise kinematic constraints, the gap between ExecVLA and other models gets substantially larger. That suggests a real advantage when you need that fine-grained control.
Dev: And they show that they can perform much more flexible and interpretable operations directly from those instructions, instead of just producing rigid motions dictated by the overall task goal. This means the robot is executing actions in a way that respects the instruction-level kinematic specifications.
Taro: Speaking of flexibility, I wonder about interpretability. If this system is so good at following these constraints, how do we know *why* it followed them? Can we probe its reasoning?
Rosa: They introduce an intervention-based analysis where they replace tokens with mismatched ones to see the effect. They found that swapping these tokens causes a substantial drop in the kinematics-following success rates, which proves those bi-level reasoning tokens play a causal role in grounding those kinematic constraints into the action generation process.
Dev: That’s solid evidence for consistency. From an engineering standpoint, they also mentioned that the inference speed doesn't add a significant overhead compared to simpler VQ-VAE models or diffusion methods, which is important for real-time applications.
Taro: So, to sum up what we have on ExecVLA, it’s about explicitly separating the abstract goal from the concrete motion parameters using this bi-level approach and linking them through reasoning tokens to ensure the robot adheres to those fine constraints. What’s next for this research direction?
Rosa: They conclude that modeling kinematics as a first-class component is essential when you need kinematics-rich tasks, and they say this bi-level formulation is naturally extensible for whole-body manipulation with more complex dependencies down the line. It’s a solid foundation to build on.
Dev: Yeah, the implication here is that for any robot that needs to manipulate objects with high precision—like those wine bottle examples mentioned in the dataset descriptions—this separation of concerns between goal and motion is key. We’re moving toward models that aren't just guessing the next step based on what they *think* they should do, but following explicit kinematic guidance.
Taro: It feels like a step towards systems that can truly follow complex, multi-faceted instructions without losing track of the high-level objective. It moves VLA from just semantic understanding to actual physical execution fidelity.
Rosa: Well, we’ve covered the main points of ExecVLA today: how they use bi-level action decomposition and reasoning tokens to handle fine kinematic constraints in vision language action models. Keep an eye on their work as it moves toward those more complex, whole-body manipulation scenarios they mentioned.
Dev: We’ll be back next time to discuss something completely different, but for now, that’s our take on ExecVLA.
The paper's summary: Rosa: So, ExecVLA basically takes what we know about vision language action models and makes them actually follow really specific motion instructions, not just vague goals.
Dev: It’s like they’ve built a system that separates the big picture objective from the tiny details of how the robot has to move to get there.
Rosa: Exactly. The paper introduces this bi-level structure which breaks down actions into two different types of representations—one for the general task goal, and one for all those fine motion parameters like exact distance or velocity.
Dev: That’s the core idea, right? They use a two-stage process to train these layers so that the fine-grained part actually learns those subtle motions really well.
Rosa: And they link this internal action representation directly to language through bi-level reasoning tokens. So you’ve got coarse reasoning for the main goal and fine reasoning for those specific kinematic anchors.
Dev: That whole system uses mutual information regularization to make sure what the robot is actually doing matches what it’s thinking in that text. It forces alignment between the language and the action execution path.
Rosa: The big result is that they show state-of-the-art performance specifically when judging how well the robot followed those precise kinematic rules, even though general goal success rates are similar to other methods.
Dev: It means if you need a robot to do something like control a wine bottle and have it face a specific orientation, this setup is much better at getting that right than standard models.
Rosa: But they also show that you can check *why* the robot did something by looking at those reasoning tokens. If you swap out the tokens with mismatched ones, the robot’s ability to follow those kinematic constraints drops significantly.
Dev: That intervention analysis is interesting because it proves those intermediate reasoning steps aren't just noise; they are causally linked to following the motion instructions.
Rosa: So, this isn't just about being smarter at understanding language; it’s about adding a layer that enforces physical reality on top of the language understanding.
Dev: It moves the robot from guessing the next step based on what it thinks is right to actually executing a plan that respects those strict kinematic boundaries.
Rosa: And they say this bi-level formulation isn't just a neat trick for one task; it can be easily adapted for whole-body manipulation where there are even more complex physical dependencies.
Dev: So, the takeaway is that if you’re building something that requires precise, instruction-level movement—like manipulating objects with specific orientations—you need to model those kinematics as a first-class component from the start.
Rosa: We’ll be looking at how this holds up in the real world and whether it stays fast enough for live control loops next time.
The paper's improvements: Tom: So, we're looking at how they suggest making ExecVLA even better for real deployment. Rosa, what’s their take on extending this beyond just following specific instructions?
Rosa: They focus a lot on how they can handle more complex physical dependencies. The authors point out that the bi-level structure is naturally extensible to whole-body manipulation, which means handling things where multiple parts of the robot need to move together with different constraints.
Dev: That makes sense for hardware, but from an engineering standpoint, how much does adding more layers or more complex physical interactions impact that loop rate we talked about? I want to know if it introduces too much latency.
Rosa: They’re trying to keep the inference speed manageable though. The authors show that even with these richer dependencies, they don't see a massive overhead compared to simpler single-level models.
Dev: That’s good news for deployment then. But what about robustness? When things get messy in the field, does this bi-level approach still hold up when the visual input is degraded or noisy?
Taro: The focus there is on using those explicit reasoning tokens to help the system recover when it gets confused by the environment. It’s about that alignment between what the language says and what it actually sees in real-time.
Rosa: They introduce an intervention-based analysis method to prove this robustness. You can swap out those reasoning tokens with bad ones and you see a sharp drop in the robot’s ability to follow its motion instructions.
Dev: That's powerful evidence, showing that those tokens actually have a causal role in grounding the movement constraints, not just being decorative text. That’s what we need for reliable control systems.
Taro: It means if the AI gets confused by an unexpected obstacle, it can use that reasoning layer to figure out which kinematic parameters are still valid and keep moving correctly towards the goal.
Rosa: Exactly. This is about improving interpretability too. Researchers can look at those tokens and see exactly what parts of the instruction—the high-level goal or the specific velocity—are driving the action.
Dev: That’s a huge win for debugging failures, because instead of just seeing a failed trajectory, you can pinpoint whether it was a goal misunderstanding or a kinematic calculation error.
Taro: It moves autonomy from just "doing what looks right" to "doing what the instruction specifically requires at this exact moment."
Rosa: So this framework helps robots understand and execute fine-grained motion instructions by cleanly separating the general task objective from the precise physical requirements.
Dev: It sets a solid foundation for making these models work on more complicated tasks, provided we can keep that inference speed in check.
Conclusion: Rosa: So we're wrapping up ExecVLA, which is all about using a bi-level action representation to let AI robots follow really precise motion instructions while keeping their overall goal steady.
Dev: It’s essentially taking the coarse task goal and separating it from the fine kinematic details so the system doesn't get overwhelmed by too much complexity at once.
Rosa: The implication is that we can finally get robots to do things that require genuine physical precision, like controlling an object to face a very specific orientation on a shelf.
Dev: But we have to be careful about the loop rate here. While they say the inference speed isn't too bad, real-time execution under load is still going to be the real test for this bi-level setup.
Taro: I think what really stands out is that because of those reasoning tokens, if things go sideways in a messy environment, the robot has a better way to figure out which kinematic constraints are actually still possible.
Rosa: That's right, Taro. It’s about making the AI more robust when it’s not in a perfect simulation but actually interacting with the physical world.
Dev: We saw that intervention analysis showed those tokens really matter for grounding the action execution, which is what engineers need to see before we trust a system on a production line.
Taro: It means the autonomy isn't just guessing; it’s following explicit, annotated paths down in the latent space of the vision language model.
Rosa: That’s exactly what makes it interesting for real-world deployment, because if you can reliably follow those fine constraints, you open up a whole new class of manipulation tasks.
Dev: I think we need to keep watching how they handle those complex dependencies when the robot needs to coordinate multiple parts of its body simultaneously.
Rosa: We will definitely do that next time, looking at how this concept scales up for whole-body tasks and more complicated physical setups.
cs.RO, cs.AI
Submitted: 2026-03-18
Updated: 2026-10-08
Code: https://github.com/Stanford-ILIAD/openvla-mini
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 89/100
The gist: The gist The KineVLA framework introduces a kinematics-rich vision-language-action (VLA) task that explicitly decouples goal-level invariance from kinematics-level variability through a bi-level
Key concepts
- Bi-Level Action Representation
- This method breaks down robot actions into two separate latent spaces: one for general task goals (coarse) and another for specific motion details like direction and velocity (fine). These two levels work together to capture both the high-level intent and the exact physical movements required by an instruction.
- Bi-Level Reasoning Tokens
- These are special tokens used during reasoning that separate the task description into two parts. The first token describes the general goal in plain language, while the second token specifies crucial kinematic details like anchor points or precise movement parameters. This structure helps connect what is said in text to how the robot should physically move.
- Goal-Level Invariance vs. Kinematics Variability
- The framework separates task goals from motion execution details. The goal level remains consistent regardless of the specific path taken, ensuring the overall objective is met. Conversely, the kinematics level allows for significant variation in how that goal is achieved, enabling fine-grained control over movement.
- Mutual Information Regularization
- This technique ensures consistency between what the model reasons about and what it actually executes. It mathematically forces a strong link between the generated reasoning tokens (text) and the resulting action, making the robot's behavior more predictable and accurate when following complex instructions.
Terminology
Summary
The gist The KineVLA framework introduces a kinematics-rich vision-language-action (VLA) task that explicitly decouples goal-level invariance from kinematics-level variability through a bi-level action representation and bi-level reasoning tokens to serve as explicit, supervised intermediate variables that align language and action.
KineVLA Framework Overview
KineVLA is a vision-language-action framework designed to address the challenge where task goals remain invariant while execution trajectories must adapt to instruction-level kinematic specifications. This framework is built upon the foundation of OpenVLA. The core innovation lies in constructing a bi-level action representation that decomposes robot actions into two complementary latent spaces: a goal level codebook that captures semantic goals and task goal, and a kinematics level codebook that encodes precise motion parameters such as direction, distance, and velocity.
Bi-Level Action Representation
The bi-level vector quantized action representation is implemented via a Bi-Level Residual VQ-VQE. This representation factors actions into two discrete spaces backed by two codebooks of identical capacity: a goal (coarse-grained) level and a kinematics (fine-grained) level. The training involves a two-stage schedule: Stage I pretraining with coarse-grained actions AGt:t+H, and Stage II finetuning with kinematics-aware actions AKt:t+H. This design encourages the fine-grained codebook to capture kinematicssensitive action representations, enabling it to model subtle action variations more effectively.
Bi-Level Generation and Reasoning Tokens
The bi-level generation paradigm uses bi-level reasoning tokens to align internal representations with language parsing. The coarse reasoning text expresses the general task goal in natural language, while the fine reasoning text specifies key kinematics parameters and anchor points. To ensure consistency between reasoning and control, a mutual information regularization scheme is introduced to maximize conditional mutual information I(Reasoning; Action C).
Datasets and Experimental Results
To support this task, three Kinematics-Rich Datasets were constructed spanning both simulation and real-world robotic platforms. These datasets feature instruction-level kinematic variations and bi-level annotations, covering diverse tabletop organization and object manipulation scenarios. Extensive experiments show that KineVLA achieves stateof-the-art performance on kinematics-aware benchmarks. Quantivative Performance Analysis shows that while all methods achieve comparable performance in terms of goal success rate across both simulation and realworld benchmarks, when evaluated on kinematics success rate, the performance gap becomes substantially larger.
Ablation Study Findings
The ablation results confirm the effectiveness and complementarity of the proposed modules. Starting from Baseline + Bi-Rep + Bi-Rea + MI, KineVLA achieves the best results of 76.5% on LIBERO-Goal-Relabeled and 70.4% on Kine-LIBERO. This demonstrates that aligning textual reasoning with action execution leads to more consistent and robust behavior. The inference speed analysis indicates that KineVLA does not introduce a significant inference overhead compared to single-level VQ-VAE models and diffusion-based methods.
Conclusion
In this work, we introduce KineVLA, a kinematics-rich vision–language–action framework that enables robots to understand and execute fine-grained motion instructions by explicitly disentangling kinematic sensitivity from goal invariance. This design allows a single goal to be realized through actions with varying kinematic granularity, improving both flexibility and interpretability. The proposed bi-level formulation is naturally extensible to whole-body manipulation with more complex kinematic dependencies. The results confirm that modeling kinematics as a first-class component is essential for kinematics-rich tasks.
Improvements for AI systems
- Bold Header: Bi-level Action Representation Implementation
KineVLA decouples goals from kinematics by "decomposing robot actions into two complementary latent spaces: a goal level codebook that captures semantic goals and task goal, and a kinematics-level codebook that encodes precise motion parameters such as direction, distance, and velocity. This allows the model to learn
highly precise motion representations" by separating low-frequency structure from high-frequency corrections.
- Bold Header: Bi-level Reasoning Token Generation
The system employs a bi-level chain-of-thought (CoT)-style generation paradigm, in which the model jointly generates textual reasoning (Reasoning) and discrete action tokens (Action).
This is intended to serve as explicit, supervised intermediate variables that align language and action,
leading to better grounding of kinematic constraints.
- Bold Header: Mutual Information Regularization for Coherence
The introduction of mutual information regularization, maximizing I(Reasoning; Action C),
is designed to ensure consistency between the generated reasoning and the control trajectory. This prevents semantic disconnection by ensuring textual reasoning may not faithfully reflect the underlying motor intention
is addressed through alignment.
- Bold Header: Fine-grained Kinematics-Aware Control
The improved system can execute fine-grained and personalized manipulation
because it can process instruction-level kinematic specifications
such as direction, trajectory, orientation, and relative displacement.
This enables the robot to perform actions like control the wine bottle to face a specific orientation on the cabinet,
which existing models cannot achieve.
- Bold Header: Interpretability via Reasoning Analysis
The system provides enhanced interpretability by allowing researchers to assess necessity through intervention-based analysis
where replacing tokens with mismatched ones causes a substantial drop in kinematics-following success rates.
This proves that bi-level reasoning tokens play a causal role in grounding kinematic constraints into action generation.
Abstract
We study fine-grained execution-constraint following in vision-language-action (VLA) models. Given an invariant task goal, the policy must follow instruction-specified execution constraints, including interaction targets, motion patterns, spatial relations, and terminal configurations. This setting exposes a limitation of goal-oriented VLAs: trajectories that complete the same task are not interchangeable when the instruction specifies how the task must be executed. We propose ExecVLA, a framework that separates a goal-oriented component from an execution-specific component through a bi-level action representation and supervised bi-level reasoning tokens. We further introduce explicit goal-invariance and execution-predictability objectives so that the goal-level representation remains stable across executions of the same goal, while the execution-level representation retains the constraints that distinguish those executions. We construct execution-constraint-following datasets in simulation and on a Realman-75 robot, with goal and fine-grained reasoning annotations. Experiments on LIBERO and the real robot show improved goal completion and, more importantly, substantially more reliable adherence to instruction-specified execution constraints.
Sources
- VLA-Touch: Enhancing Vision-Language-Action Models with Dual-Level Tactile Feedback
- $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
- RT-1: Robotics Transformer for Real-World Control at Scale
- Tactile-VLA: Unlocking Vision-Language-Action Model's Physical Knowledge for Tactile Generalization
- $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
- Behavior Generation with Latent Actions
- HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model
- RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation
- ManiSkill: Generalizable Manipulation Skill Benchmark with Large-Scale Demonstrations
- Representation Learning with Contrastive Predictive Coding
- FAST: Efficient Action Tokenization for Vision-Language-Action Models
- VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers
- FreeTacMan: Robot-free Visuo-Tactile Data Collection System for Contact-rich Manipulation
- Qwen3 Technical Report
- ForceVLA: Enhancing VLA Models with a Force-aware MoE for Contact-rich Manipulation
- DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving