When Instructions Retrieve Trajectories: Diagnosing and Mitigating Generalization Failures in VLA Models
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: I'm Rosa, and with me are Dev and Taro, guest researcher.
Dev: Today's paper: "When Instructions Retrieve Trajectories".
Rosa: Vision-language-action (VLA) models can exceed 90% success on in-distribution tasks and withstand nuisance changes that preserve the required action, yet fail under counterfactual changes that demand a different action.
Dev: First, who's behind it and why it matters.
Paper summary: Rosa: Welcome back to the show, everyone. We're talking about a really interesting paper from arXiv titled "When Instructions Retrieve Trajectories: Diagnosing and Mitigating Generalization Failures in VLA Models." The main point of this paper is that vision-language-action models, even when they hit over ninety percent success on tasks they’ve seen before, can completely fail when the instructions ask for something different than what they are used to. They call this failure instruction-action binding.
Dev: That sounds like a problem where the AI gets stuck in a rut because it's listening to both the vision and the language but isn't properly merging them to pick the right move. I wonder if this is something we see often in our control loops, Rosa?
Taro: From an autonomy standpoint, this binding suggests that when things go wrong in a real environment, the AI doesn't just shut down; it often defaults to a behavior it's already seen before instead of figuring out the new requirement.
Rosa: Exactly what Taro is saying. The paper claims that this binding happens because the instruction cues familiar trajectory families, and the visual feedback then just adjusts how that familiar family runs, instead of changing which family is used at all. This means instruction-keyed solutions can fit the supervision without actually learning how to change actions when something new comes up in a scene.
Dev: That leads right into my concern about latency and failure modes, Rosa. If the policy is just selecting a familiar trajectory family based on the instruction key, what happens when the visual feedback strongly contradicts that choice? Does that lead to immediate instability in our execution loop?
Taro: Well, if we look at how these failed rollouts behave in behavioral analyses of fine-tuned policies like π0 point 5 and GR00T-N1 point 7, they often either stick with the original behavior or switch over to another task they've demonstrated before.
Rosa: That switch is really telling because it shows that the language isn't just being ignored by the system; instead, it’s selecting a familiar path, and then the visual feedback just adapts that path’s execution. This suggests we need to look deeper into what internal state is driving those choices.
Dev: I see why you bring up the internal state; from a control engineering view, if the system is selecting a trajectory family that doesn't satisfy the current instruction, that implies a breakdown in how the instruction and vision are being combined at that decision point.
Taro: The paper models this by contrasting two policies: one grounded policy that resolves the instruction's reference within the scene through pi ground(a s,) = pi (a g(s, r)) and an idealized instruction-keyed lookup policy that selects actions based on "memorized keys" according to pi lookup(a s,) = pi (a k).
Paper summary: Rosa: That contrast helps explain why they see this gap; the failure occurs when local visual feedback can still be active even if the selected trajectory family doesn't meet the current instruction and scene requirements. It’s about a disconnect between what's being requested and what's actually happening visually.
Dev: And they quantify this by looking at changes in the instruction key, denoted as k, versus a change in the required action, a, showing instances where there is "same key, different required action," which they label Swap zero one. This gives us a concrete way to spot that binding happening in the data.
Taro: That quantification is useful because it moves beyond just saying a failure happened; it tells us *how* the system is failing at a mechanistic level, which helps us design better safeguards for when things misbehave in the real world.
Rosa: Moving on to how they try to fix this, they introduce something called Equivariant Counterfactual Training or ECT. This approach tries to bridge that supervision gap by supplying valid demonstrations where the same instruction demands different actions in scenes that are easily distinguishable.
Dev: How does the ECT loss actually work operationally? I need to know if this is adding significant computational overhead or introducing new forms of instability into our training pipeline.
Taro: The ECT loss is defined as L ECT(theta) = E s about E
L theta(s,, a) + L theta(s',, a'): , where both the standard and the counterfactual counterparts contribute to the same parameter update.
Rosa: It forces the model to learn that these different scenes, even with the same instruction, require different actions by training on those paired demonstrations together. They construct action-valid counterfactual pairs using operators like M s(s) for scene change and M a(a) for path transformation to ensure the changed scene genuinely requires a different action under the same instruction.
Dev: Building on that, the results they show are quite strong across different setups, like LIBERO-PRO, CALVIN, and even physical UR5e demonstrations. In the LIBERO-PRO comparisons, full ECT raised pi zero point five ’s mean position-swap success from thirty-six percent up to fifty-nine percent.
Taro: That jump in performance is substantial because it shows that this training method actually addresses the core problem of generalization, not just boosting a specific metric on one dataset. It suggests that providing these scene-dependent alternatives during training really helps the model internalize the instruction-action relationship more deeply.
Rosa: I think what's important here for us is that they show ECT data and loss are complementary; the ECT data creates those missing scene-dependent alternatives, while the ECT loss makes their correspondence explicit during training. The study also found that constructed data actually improve over standard demonstrations in some cases, like in CALVIN where the ECT loss alone raised five-task completion from fifty-eight percent to seventy-six percent.
Paper summary: Dev: That’s good news regarding performance gains, but I have to mention the limitations they point out; they note that Swap and Task remain below identification level even after ECT in every suite. This suggests that the binding might characterize task selection rather than just a unique retrieval algorithm, which is something we need to keep in mind for our deployment plans.
Taro: The limitation they mention is important because it pushes us toward thinking about how these systems operate outside of the lab; if binding characterizes task selection, it means we need to ensure the system can handle novel tasks that are structurally different from what it has learned, not just variations of old tasks.
Rosa: So, when we talk about the implications of "When Instructions Retrieve Trajectories: Diagnosing and Mitigating Generalization Failures in VLA Models," we're talking about moving toward systems that don't just perform well on known tasks but actually understand *why* they are performing them and can adapt when the context shifts unexpectedly.
Dev: It means that if we want to deploy these kinds of models on more complex physical robots, we need to rigorously test them against counterfactual scenarios where the instruction is valid but the visual scene demands a different outcome. The latency concerns remain, though, because running this kind of comparison adds complexity to the decision-making loop.
Taro: The bigger picture implication for autonomy is that we need explicit mechanisms to handle ambiguity when the instruction and vision conflict in a way that doesn't just lead to a generic fallback behavior. We need systems that can reason about what the instruction *means* in the current visual context, not just what trajectory family it points toward.
Rosa: It sounds like this paper is pushing us to realize that evaluation needs to go beyond simple success rates on known examples and actually probe for this kind of instruction-action binding failure in the real world.
Dev: I think the focus now should be on how we can integrate these counterfactual checks efficiently into the inference pipeline without crippling the loop rate, which is always my primary concern when looking at these types of models.
Taro: Ultimately, if we can solve this binding issue, it means VLA models become much more reliable for complex physical tasks where instructions are dynamic and environments change constantly.
Rosa: Well said. That's what we have covered regarding the core findings and what the authors are proposing with ECT in "When Instructions Retrieve Trajectories: Diagnosing and Mitigating Generalization Failures in VLA Models."
Conclusion: Rosa: So, we've seen how these VLA models struggle when instructions conflict with what they see in real-time, and now we're wrapping up this discussion on "When Instructions Retrieve Trajectories: Diagnosing and Mitigating Generalization Failures in VLA Models."
Dev: That paper really zeroes in on that instruction-action binding failure mode, where the AI picks a familiar path instead of the one it was actually told to take. The authors are showing us how this happens by contrasting how a grounded policy works versus an idealized lookup policy.
Taro: I think what's striking is their breakdown of *why* this happens—it isn't just that the language is ignored; it’s because the instruction selects a familiar trajectory family and the vision just adapts that execution. That distinction is key for understanding system behavior when things go wrong in an autonomous setting.
Rosa: And their proposed fix, Equivariant Counterfactual Training, ECT, seems like a clever way to force the model to learn those scene-dependent alternatives by training on paired demonstrations where the same instruction requires different actions in different scenes.
Dev: From my side as a control engineer, I'm interested in how robust this method is; if we are running these complex counterfactual checks, I have to worry about the loop rate and whether this adds too much latency to our decision-making process during actual deployment.
Taro: That’s a fair concern, Dev, but the results on physical platforms like the UR5e show that it can dramatically improve success rates when dealing with unseen scenarios under fixed demonstration budgets. That suggests a potential path toward better handling of real-world unpredictability.
Rosa: It really makes you think about how these systems will perform long-term outside of a controlled lab setting; will this improved generalization hold up when faced with genuinely novel physical situations over extended operation?
Dev: I'd say the paper confirms that while ECT helps significantly, they also noted that Swap and Task metrics stay below identification level even after training, which suggests the binding might be tied more to task selection than a single retrieval algorithm.
Taro: Exactly, so this points toward needing systems that can reason about the instruction’s meaning in context rather than just relying on memorized trajectories for every scenario. This has big implications for how we design truly flexible autonomy.
Hung-Jen Chen, Yu-Hsun Hou, *Yan-Hong Chen, *Yan-Fu Chen, Binghua Cai, Min Sun, Chun-Yi Lee
National Tsing Hua University
cs.RO, cs.LG
Submitted: 2026-09-30
Updated: 2026-09-30
Code: https://github.com/Zxy-MLlab/LIBERO-PRO
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 83/100
The gist: Vision-language-action (VLA) models can exceed 90% success on in-distribution tasks and withstand nuisance changes that preserve the required action, yet fail under counterfactual changes that demand
Key concepts
- Instruction-Action Binding
- This is a failure mode where a VLA policy responds to both language and vision but fails to combine them correctly. Instead of selecting the action required by the instruction, it defaults to familiar trajectory families, even when the visual feedback suggests a different path is needed for the current scene.
- Nuisance vs. Counterfactual Perturbations
- The paper distinguishes between two types of changes: nuisance perturbations leave familiar behavior valid, while counterfactual perturbations demand a new action. Failed rollouts often retain source behavior or switch to another demonstrated task, indicating the model is not learning how the required action should change across scenes.
- Equivariant Counterfactual Training (ECT)
- ECT is a training method that addresses the supervision gap by creating pairs of data: valid demonstrations where the same instruction requires different actions in distinct scenes. The loss function trains both counterparts simultaneously, ensuring the model learns to distinguish between these scene-dependent action requirements.
Terminology
Summary
Vision-language-action (VLA) models can exceed 90% success on in-distribution tasks and withstand nuisance changes that preserve the required action, yet fail under counterfactual changes that demand a different action.
The gist
Instruction-action binding is a failure mode where a policy responds to both language and vision yet fails to combine them to select the action the task requires, often by selecting familiar trajectory families instead of the instruction-conditioned one.
Diagnosis of Instruction-Action Binding
The paper identifies this failure by distinguishing between nuisance perturbations, which leave familiar behavior valid, and counterfactual perturbations, which demand a new action. Behavioral analyses reveal that failed rollouts often retain the source behavior or switch to another demonstrated task.
This suggests that language is not simply ignored; instead, The instruction selects a familiar trajectory family, and visual feedback still adapts that family’s execution,
leading to an issue where instruction-keyed solutions can therefore fit the supervision without learning how the required action changes across scenes.
Mechanistic Framework
The failure is modeled by contrasting two policies: a grounded policy that resolves the instruction's referent in the scene through πground(a s, l) = π (a g(s, r(l)))
and an idealized instruction-keyed lookup policy that selects among demonstrated trajectory families according to πlookup(a s, l) = π (a k(l))
. The binding occurs when local visual feedback can remain active even when the selected trajectory family does not satisfy the current instruction and scene.
This is further quantified by distinguishing between a change in the selected instruction key, denoted as ∆k,
and a change in the required action, denoted as ∆a,
where a swap cell shows same key, different required action
(Swap 0 1).
Equivariant Counterfactual Training (ECT)
To address this supervision gap, the paper introduces Equivariant Counterfactual Training (ECT), which acts at two levels. ECT data supply valid demonstrations in which the same instruction requires different actions in distinguishable scenes,
while the ECT loss trains each demonstration with its counterpart in the same update. This is operationalized by constructing action-valid counterfactual pairs using transformation operators like Ms(s)
for scene change and Ma(a)
for path transformation, ensuring that the changed scene must require a different action under the same instruction.
The ECT loss is defined as LECT(θ) = Ee∼E [Lθ(s, l, a) + Lθ(s′, l, a′)], where both counterparts contribute to the same parameter update.
Experimental Validation and Results
The efficacy of ECT is demonstrated across multiple suites (LIBERO-PRO), including LIBERO-PRO, CALVIN, and physical UR5e demonstrations. In controlled comparisons on LIBERO-PRO, full ECT raises π0.5’s mean position-swap success from 36% to 59%.
On the real UR5e under a fixed demonstration budget, full ECT raises unseen-position success from 8% to 88%.
The analysis shows that ECT data and the ECT loss are complementary: ECT data create missing scene-dependent alternatives, while the ECT loss makes their correspondence explicit during training.
Furthermore, the study confirms that Constructed data also improve over Standard,
and in CALVIN, the ECT loss alone raises five-task completion from 58% to 76%.
Key Contributions
The central scientific contributions are threefold: (i) identifying instruction-action binding in two VLA architectures through behavioral probes and internal-state interventions; (ii) explaining how narrow conditional action support permits instruction-keyed solutions, motivating sameinstruction alternatives that require different actions and are distinguishable from the scene
; and (iii) introducing ECT data and the ECT loss to operationalize this principle. The results validate that Counterfactual evaluation is therefore essential for distinguishing genuine task-level grounding from policies that remain robust while selecting familiar but incorrect behaviors.
Limitations
The paper notes that Swap and Task remain below ID after ECT in every suite,
suggesting binding characterizes task selection rather than a unique retrieval algorithm. Additionally, the study emphasizes that ECT requires action-valid, scene-dependent alternatives,
and generalist use without such fine-tuning remains untested. The paired sampler's optimization mechanism is noted as being open for further investigation.
Improvements for AI systems
Here are the specific improvements that can be made to existing Vision-Language-Action (VLA) models, based on the findings of this paper:
-
Improved Counterfactual Robustness via Equivariant Counterfactual Training (ECT):
-
Enhanced Instruction-Action Binding Diagnosis and Mitigation:
-
Scene-Dependent Action Selection Capability:
- Improved Counterfactual Robustness via Equivariant Counterfactual Training (ECT):
The system can be trained using the ECT loss, which simultaneously trains a demonstration with its counterpart
in a different scene that requires the same instruction but demands a different action.
-
This specifically addresses failures under counterfactual changes (like swapping object positions or changing goals).
-
By training on these paired alternatives together, the model learns to distinguish between actions based on the scene context, rather than relying solely on memorized trajectories.
- Enhanced Instruction-Action Binding Diagnosis and Mitigation:
The system can be diagnosed by monitoring internal states (like prefix-KV representations) and using interventions to verify if it is following a familiar trajectory family or making a task-conditioned selection.
-
If the model selects an action based only on the instruction key (instruction-keyed solution), the system can be prompted with scene cues to force it to select a different, scene-dependent action.
-
The system can be explicitly trained to resolve this binding by providing supervision that forces it to choose between distinguishable scenes requiring different actions under the same instruction.
- Scene-Dependent Action Selection Capability:
The improved system can select the correct action based on the current visual scene, even when provided with a familiar instruction.
-
The system will no longer default to executing a familiar behavior simply because it matches the source demonstration's trajectory family (i.e., it will not suffer from
instruction-action binding
). -
This allows for more flexible and generalizable manipulation, as the policy uses both language and vision modalities together to select the action appropriate for the current scene/instruction combination, rather than just retrieving a memorized trajectory.
Sources
- GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
- $\pi_\texttt{RL}$: Online RL Fine-tuning for Flow-based Vision-Language-Action Models
- ProGAL-VLA: Grounded Alignment through Prospective Reasoning in Vision-Language-Action Models
- When Vision Overrides Language: Evaluating and Mitigating Counterfactual Failures in VLAs
- LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models
- CAST: Counterfactual Labels Improve Instruction Following in Vision-Language-Action Models
- Mechanistic interpretability for steering vision-language-action models
- OTTER: A Vision-Language-Action Model with Text-Aware Visual Feature Extraction
- $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
- Don't Blind Your VLA: Aligning Visual Representations for OOD Generalization
- OpenVLA: An Open-Source Vision-Language-Action Model
- LangForce: Bayesian Decomposition of Vision Language Action Models via Latent Action Queries
- Flow Matching for Generative Modeling
- Unrolled Generative Adversarial Networks
- Seeing to Act, Prompting to Specify: A Bayesian Factorization of Vision Language Action Policy
- LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving