Same Scene, Different Task: Skill Alignment for Compositional Generalization in VLAs
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: Today's paper: "Same Scene, Different Task".
Dev: Fine-tuned Vision-Language-Action (VLA) models often struggle to generalize to skill combinations absent from their fine-tuning demonstrations, even when every constituent skill has been demonstrated.
Rosa: First, who's behind it and why it matters.
Title and authors: Rosa: So we're starting with "Same Scene, Different Task: Skill Alignment for Compositional Generalization in VLAs," and I'm really interested in what this paper is tackling because compositional generalization is such a big hurdle for robots right now.
Dev: Yeah, it sounds like they are diving into that exact problem where models get stuck on combinations they haven't explicitly seen during training. What’s the core idea behind their approach to solving that?
Taro: Basically, the paper points out that when a task is made of smaller skills, like picking and placing, and you only show them those skills separately in demonstrations, the AI struggles when you ask it to do a new combination like picking one object and placing it on a different surface.
Rosa: Exactly. The title suggests they are trying to fix this by aligning skill representations so the model doesn't rely too heavily on just looking at what's in front of it instead of paying attention to the specific instruction for that moment.
Dev: I saw their summary mentions that they use existing demonstrations of constituent skills to train VLAs to execute skill combinations absent from the original demonstrations without having to collect new data. That sounds like a smart way around the data collection bottleneck we usually face in these VLA systems.
Taro: That’s interesting because it avoids the huge cost of generating entirely new trajectories for every possible combination, which is what some other approaches have tried to do by creating synthetic demonstrations.
Rosa: Right, and their methodology seems to hinge on using skill representations that can be reused across different executions of the same basic skill, while also allowing action predictions to change based on the observation. It sounds like they are trying to separate *what* skill is being done from *where* exactly it's happening in the scene.
Dev: And I noticed they introduce these learnable query tokens, Qskill and Qstate, where the skill queries look at both image and text tokens to build a skill representation, while state queries only look at the image for a state representation. That sounds like a clever way to give the model different ways to interpret visual information.
Taro: I think that separation is key; if we can isolate the pure skill concept from the specific visual context, it might help when things get messy in real-world scenarios where unexpected stuff happens.
Title and authors: Rosa: And then they layer on three training objectives: a standard flow-matching loss to supervise the action expert with what they already have, an Lskill objective to make those representations reusable across different skill executions, and an Lcf for counterfactual skill-representation alignment.
Dev: That counterfactual supervision part sounds crucial; it’s not just about matching demonstrations anymore; it’s about training the representation to correctly infer the required skill even when the instruction is changed in a way that's absent from the original data.
Taro: If that Lcf objective works as they suggest, it means we can give them supervision for a new combination using an old demonstration of just one of its parts, which is exactly what we need for robustness.
Rosa: So, to sum up what we've heard about "Same Scene, Different Task: Skill Alignment for Compositional Generalization in VLAs," they are proposing a method that uses existing skill demonstrations to teach the VLA model how to handle combinations it hasn't seen before by focusing on learning reusable skill representations.
Dev: It seems like the practical application is making these models significantly better at generalizing, even when the instruction requires chaining skills in a novel way. I wonder if this level of generalization translates well when we look outside the clean simulation environment.
Taro: That’s my main concern about deployment; if it works perfectly in simulation but fails when the world misbehaves, that’s where we have to be careful. The paper doesn't explicitly detail how this handles novel physical interactions outside of their structured benchmarks.
Rosa: That brings us right to the next point: the improvements they suggest for this framework, which seem designed to make it even more robust and versatile than what was presented in the core paper.
Dev: The improvements focus on making those skill representations truly reusable across different executions of that same skill, while also ensuring they are distinct enough for different skills. It’s about managing that trade-off between reusability and specialization.
Taro: And I like the idea of using counterfactual supervision to train the representation to reflect a required skill under a changed instruction, rather than just matching the input observation. That should give it more reasoning power when faced with ambiguity.
Title and authors: Rosa: It seems they are building on their initial concept by adding mechanisms that enforce this skill identity separation and use the counterfactual pairs actively during training, which should make the system more flexible in real-world deployment situations where instructions might be phrased differently.
Dev: From an engineering standpoint, I’m focused on how this affects latency and loop rate; if these new representations add too much complexity to the inference pipeline, we could see performance hit our required real-time constraints. I'd want to see details on the computational overhead of those Qskill and Qstate queries.
Taro: That's a fair point, Dev; any extra layers in the policy architecture need to be efficient enough so they don't introduce unacceptable delays in a dynamic environment where speed matters for safety.
Rosa: So, we’ve covered the core idea of how "Same Scene, Different Task: Skill Alignment for Compositional Generalization in VLAs" uses counterfactual supervision to tackle skill combination failures by learning reusable skill representations. We're now looking at how the authors suggest pushing this framework further with specific architectural improvements.
Dev: And those improvements seem aimed squarely at solving the generalization gap while keeping the system computationally feasible, which is a tough balancing act for any VLA system.
Taro: I think if they can successfully decouple skill identity from specific observations, that opens up possibilities for more adaptable autonomy when things go wrong in unexpected ways.
Rosa: Indeed, it sounds like the paper lays a solid foundation by showing that we don't need new data for every combination if we use these clever alignment techniques to train the underlying skill understanding.
Dev: So, as we wrap up this discussion on "Same Scene, Different Task: Skill Alignment for Compositional Generalization in VLAs," it seems like they offer a concrete way to improve how models handle complex task sequencing without needing massive amounts of new demonstration data.
Taro: I think the real impact here is showing that existing demonstrations can be leveraged far more effectively than we currently do when dealing with sequential actions.
Rosa: Absolutely, and this work provides valuable insight into structuring the training supervision to encourage better compositional understanding in VLA models. We’ll take a quick pause before we look at what these findings actually mean for the wider field of robotics.
The paper's summary: Rosa: So, we've been hearing about "Same Scene, Different Task: Skill Alignment for Compositional Generalization in VLAs," and now I want to talk about what that summary actually means for us as field roboticists and engineers.
Dev: It seems the core of the paper boils down to using existing demonstrations of individual skills to train a VLA model so it can successfully execute entirely new skill combinations it hasn't seen before, all without needing new data collection.
Taro: That’s what interests me about the counterfactual supervision part; essentially, they're using what they already know about one skill to teach the model how to handle another when the instruction changes in a way that wasn't in the original training set.
Rosa: Exactly, it tackles that vision shortcut problem where models get confused by visual cues instead of following the actual sequence of operations specified in a complex task.
Dev: From an engineering standpoint, I’m thinking about how this impacts our loop rates; if these learned skill representations are too slow or complex to process, we won't get the real-time performance we need for deployment.
Taro: That's a valid concern, Dev; the paper suggests they introduce mechanisms to make those representations reusable across different executions of the same skill, which should theoretically keep the computational load manageable while boosting generalization.
Rosa: It sounds like their results show that this approach improves success on combinations we haven't demonstrated—like picking an object and placing it in a new spot—while still keeping high success on the combinations they actually showed us during fine-tuning.
Dev: The evaluation results, especially when you look at real-robot testing, are pretty impressive; they show significant jumps in performance on those undemonstrated tasks compared to standard fine-tuning methods.
Taro: I'm really excited about the implication here for autonomy; if we can reliably generalize from known skills, it means we can build systems that are much more flexible and less brittle when faced with unexpected scenarios outside of the lab.
Rosa: That’s what makes me keen to know how long this works in practice; does this generalization hold up when the environment gets messy or when things move unexpectedly?
Dev: The paper’s limitations are clear, though; they state that this formulation assumes a fixed sequence of operations and that every single skill needed for an undemonstrated combination was present in the original demonstrations.
Taro: That limitation is important; so if we have a completely novel operation or a brand new skill entirely, this specific framework won't cover it without modification.
Rosa: Well, that means we need to watch how the authors propose expanding this concept to handle those more open-ended tasks and varying instruction sequences in the future.
The paper's improvements: Rosa: So, we've discussed how "Same Scene, Different Task: Skill Alignment for Compositional Generalization in VLAs" uses counterfactual supervision to teach models about skill combinations without needing new data.
Dev: Right, and now we're looking at the suggested improvements they propose to make this approach even more robust and effective for real-world applications.
Taro: I'm looking at the points where they suggest learning a reusable skill representation that is distinct from others, which should really help when the model has to switch between different skills quickly.
Rosa: It sounds like they're focusing on how to make that skill identity separation clearer so the VLA can properly condition its actions based on what it needs to do at that specific moment.
Dev: I’m focused on the computational cost here; these additional learning objectives, like Lskill and Lcf, introduce more complexity into the training process, so I want to know how much overhead we’re looking at during inference when running these improved models.
Taro: The counterfactual supervision objective they mention is really interesting because it trains the skill representation to reflect the exact skill required by a changed instruction under fixed reference states.
Rosa: That sounds like it gives us more reasoning power; it means when an instruction is phrased in a way we haven't seen, the AI has a better chance of inferring the correct underlying action requirement.
Dev: If this improved system can successfully maintain high success on demonstrated combinations while substantially increasing success on undemonstrated ones, that’s a significant step toward deploying these agents in more complex environments.
Taro: I think this directly addresses the brittle nature of current VLA systems; instead of failing completely when an instruction shifts slightly, it should be able to adapt its behavior based on its learned skill representations.
Rosa: That opens up possibilities for much more versatile robotic assistants, capable of handling a wider variety of tasks without needing entirely new training datasets for every single variation.
Dev: We need to see if this adaptability translates into reliable, low-latency execution during actual operation, because a clever representation that takes too long to compute is just useless in a fast-paced physical world.
Taro: The paper flags a limitation where this method assumes the task sequence is fixed and every skill needed was shown in the original demonstrations; so we still need to address how it handles genuinely novel or entirely unrepresented skills.
Rosa: Exactly, so the future work needs to focus on extending this framework beyond fixed sequences and incorporating mechanisms for learning those entirely new skills dynamically.
Conclusion: Rosa: So, to wrap things up, we've seen how "Same Scene, Different Task: Skill Alignment for Compositional Generalization in VLAs" uses counterfactual supervision to train VLA models on skill combinations they haven't explicitly seen before.
Dev: It really shows a way to boost the system’s ability to generalize from learned skills without needing massive amounts of new demonstration data.
Taro: I think the real value is in how it handles uncertainty; by aligning skill representations this way, we get better reasoning when the world throws us an instruction that doesn't perfectly match our original training examples.
Rosa: And that’s exactly where we need to watch: whether this robustness holds up when we take these agents out of the lab and into truly unstructured environments for extended periods.
Dev: From my side, I’m still checking the computational overhead; if the complexity of those skill representations causes any significant latency spikes during high-speed execution, that’s a failure mode we can't ignore.
Taro: That adaptability is what makes this concept important for autonomy; if we can make systems more flexible when things misbehave, it moves us closer to truly reliable navigation and manipulation in the wild.
Rosa: It’s exciting to think about how this could help build robots that are less prone to breaking down when faced with a slightly different task sequence.
Dev: We need more data on the long-term stability of these learned skill representations under continuous operation, not just benchmark success rates in simulation or controlled settings.
Taro: So, we’re left wondering how far this technique can stretch before it hits its limits when dealing with completely new physical interactions that aren't covered by the existing skill demonstrations.
Rosa: That’s the open question for our field: what are the necessary next steps to push this framework into truly general-purpose autonomy?
Dev: I'm curious about how they plan to make these learned skills more resilient against external disturbances that might affect the state representation during operation.
Taegeun Yang, Youngju Na, Yoonki Cho, Sung-Eui Yoon
Korea Advanced Institute of Science and Technology (KAIST
cs.RO, cs.LG
Submitted: 2026-09-30
Updated: 2026-09-30
Comments: 26 pages, 5 figures. Project page: https://taegeunyang.github.io/craft/
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 82/100
The gist: Fine-tuned Vision-Language-Action (VLA) models often struggle to generalize to skill combinations absent from their fine-tuning demonstrations, even when every constituent skill has been demonstrated.
Key concepts
- Compositional Generalization
- This is the challenge where a VLA model knows how to perform individual skills but fails when those skills are put together in a new order or with different objects. The paper addresses this by training models to handle these combinations they haven't explicitly seen before.
- Counterfactual Pairs
- These are pairs created by keeping the visual observation the same but changing the instruction to request an undemonstrated skill combination. This setup helps train the model on what action to take when given a novel instruction, even though no direct action target was shown for that specific pair.
- Skill Representation ($z_{skill}$)
- This is a learned feature that captures the essence of a specific skill, independent of the current visual state. It is trained to be reusable across different executions of the same skill, allowing the model to apply knowledge about one skill to a new context effectively.
- Skill Query Tokens ($Q_{skill}$)
- These are learnable tokens added to the VLM sequence that focus on understanding the required skill. They attend to both visual and text information simultaneously to generate a 'skill representation' that guides the action prediction process.
Terminology
Summary
Fine-tuned Vision-Language-Action (VLA) models often struggle to generalize to skill combinations absent from their fine-tuning demonstrations, even when every constituent skill has been demonstrated. This work introduces CRAFT, a counterfactual supervision approach that uses existing demonstrations of constituent skills to train VLAs to execute skill combinations absent from the original demonstrations without collecting new data.
The Gist
CRAFT improves success on undemonstrated skill combinations while maintaining high performance on demonstrated ones by transferring supervision from demonstrated executions of the required skill to counterfactual pairs using skill representations that can be reused across executions of the same skill.
Problem Formulation and Motivation
Many robotic tasks are compositions of skills, where a task is defined by a fixed sequence of operations and an entity tuple specifying which object or location is used for each operation. The core challenge addressed is compositional generalization: fine-tuned VLAs often perform well on demonstrated combinations but struggle with new combinations because they use visual observations as a proxy for the instruction (a vision shortcut
), leading them to execute a demonstrated combination associated with similar observations rather than the instructed one. To mitigate this, the authors form counterfactual pairs
by holding a demonstration observation fixed while changing the instruction to specify an undemonstrated combination. The crucial difficulty is that these pairs lack corresponding demonstrated action targets because actions from those executions cannot serve as direct targets due to variations in required actions across observations for the same skill.
CRAFT Methodology
CRAFT addresses this challenge by learning skill representations that can be reused across executions of the same skill while allowing action prediction to vary with the observation. The policy is conditioned on both a skill representation
and a state representation.
Specifically, learnable query tokens, denoted as Qskill and Qstate, are added to the VLM token sequence. The skill queries attend to both image and text tokens to produce the skill representation (z skill), whereas the state queries attend only to image tokens to produce the state representation (z state). The action prediction is then conditioned on both: π(a o, l) = pθ(a z skill, zstate).
The training involves three complementary objectives combined in a total objective function:
-
The standard flow-matching loss (Lfm), which supervises the action expert using demonstrated action chunks.
-
Lskill, which encourages skill representations to be reusable across same-skill executions by contrasting prediction errors while holding the state representation and target velocity fixed. This is achieved by comparing predictions from different skill representations using a contrastive objective (Lskill).
-
Lcf (Counterfactual Skill-Representation Alignment), which trains the counterfactual pair’s skill representation to reflect the required skill under the changed instruction by conditioning action prediction on that new skill and supervising it with a demonstrated execution of that required skill.
Evaluation and Results
CRAFT was evaluated across three VLA models (π0, π0.5, and GR00T N1.7) on two compositional benchmarks: PICK-PLACE and PICK-PLACE-PRESS in simulation, with additional evaluation on a real robot (π0.5). The results show that CRAFT substantially improves success on undemonstrated skill combinations while maintaining high performance on demonstrated ones across all models and benchmarks. For instance, it achieves 84.2% success on undemonstrated combinations for the PICK-PLACE benchmark with the π0.5 model, compared to 62.6% for standard fine-tuning (FT). Furthermore, in real-robot evaluation, CRAFT achieved a success rate of 43/60 (71.7%) on undemonstrated PICK-PLACE combinations, outperforming FT (Full) at 9/60 (15%). The analysis of skill representations confirms that Lskill encourages representations to be reusable across same-skill executions,
and the counterfactual supervision successfully trains the skill representation to reflect the currently required skill under ˜li in action prediction.
Key Contributions
The paper makes several key contributions:
-
Proposing CRAFT, a counterfactual supervision approach that uses existing demonstrations of constituent skills to train VLAs to execute skill combinations absent from the original demonstrations without collecting demonstrations for those combinations.
-
Introducing two compositional benchmarks for two- and three-skill tasks, with demonstration datasets that cover every constituent skill but only a subset of skill combinations.
-
Demonstrating that across π-family and GR00T VLA models, CRAFT improves success on skill combinations absent from the original demonstrations while maintaining high success on demonstrated combinations; it also validates CRAFT with π0.5 on a real robot.
Limitations
The formulation assumes a fixed operation sequence and that every constituent skill in an undemonstrated combination appears in the demonstrations. Tasks requiring new skills or operation sequences that vary with the instruction are outside the scope of this formulation.
Improvements for AI systems
Based on the provided scientific paper, here are specific improvements that can be made to current Vision-Language-Action (VLA) systems by implementing the CRAFT framework, along with a description of what these improved systems can achieve:
-
The core improvement is introducing the Counterfactual Skill-Representation Alignment for Fine-Tuning (CRAFT) mechanism. This directly addresses the
vision shortcut
failure mode where models rely on visual proxies instead of following explicit instructions when new skill combinations are required. -
The improved system can perform compositional generalization on novel, unseen skill combinations that were not present in the original fine-tuning demonstrations, even if every individual constituent skill (e.g., pick and place) has been demonstrated multiple times across different pairings.
-
The system will utilize two learned representations: a reusable Skill Representation and a state representation derived from visual observation, both conditioned on the instruction.
-
The improved system can decouple skill identity from specific observations during action prediction by training the Skill Representation to be highly reusable across executions of the same skill, while simultaneously ensuring that different skills map to distinct representations.
-
During fine-tuning, the system will employ a counterfactual supervision strategy (using demonstrated executions of the required skill as targets) instead of requiring demonstrations for every specific combination.
-
The system can be trained effectively using only existing demonstrations, without needing to collect new, expensive demonstration data for every possible skill combination.
-
The system will use a composite loss function: a standard flow-matching loss (for demonstrated actions) combined with two auxiliary objectives:
-
A contrastive objective (Lskill) that encourages the Skill Representation to be reusable across same-skill executions and distinguishable from other skills in action prediction.
-
A counterfactual objective (Lcf) that trains the Skill Representation to reflect the skill required by a counterfactual instruction, using fixed reference states and targets from demonstrations of that required skill.
-
The improved system will exhibit better responsiveness to instruction changes during execution:
-
When an instruction change requires switching to a new skill, the policy will successfully execute that new skill from a state reached while executing a different skill (i.e., it learns to reuse learned representations across executions).
-
The system can maintain high success rates on demonstrated combinations while simultaneously achieving substantially higher success rates on undemonstrated combinations across various VLA models and benchmarks (as shown in Table 1).
-
This compositional generalization capability is validated both in simulation and when deployed on a real robot, suggesting a robust improvement over standard fine-tuning methods that often fail to bridge the gap between demonstrated and novel skill combinations.
Abstract
Vision-language-action (VLA) models often struggle to generalize to skill combinations absent from their fine-tuning demonstrations, even when every constituent skill has been demonstrated. We focus on a vision shortcut as one failure mode: during fine-tuning, visual observations can serve as a proxy for the instruction, so a policy may execute a demonstrated combination associated with similar observations rather than the instructed combination. This motivates training with counterfactual pairs formed by holding a demonstration observation fixed while changing the instruction to specify an undemonstrated combination. These pairs, however, lack corresponding demonstrated action targets. Crucially, the currently required skill has already been demonstrated, but actions from those executions cannot serve as direct targets because the same skill can require different actions across observations. We propose CRAFT, which transfers supervision from demonstrated executions of the required skill to counterfactual pairs using skill representations that can be reused across executions of the same skill. Across three VLA models and two simulation benchmarks, CRAFT improves success on undemonstrated combinations while maintaining high success on demonstrated ones; it also improves compositional generalization on a real robot. Project website: https://taegeunyang.github.io/craft/
Sources
- Do As I Can, Not As I Say: Grounding Language in Robotic Affordances
- GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
- $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
- $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
- ReSteer: Quantifying and Refining the Steerability of Multitask Robot Policies
- When Vision Overrides Language: Evaluating and Mitigating Counterfactual Failures in VLAs
- Efficient Data Collection for Robotic Manipulation via Compositional Generalization
- CAST: Counterfactual Labels Improve Instruction Following in Vision-Language-Action Models
- Enhancing Generalization in Vision-Language-Action Models by Preserving Pretrained Representations
- PriorVLA: Prior-Preserving Adaptation for Vision-Language-Action Models
- LoRA: Low-Rank Adaptation of Large Language Models
- OpenVLA: An Open-Source Vision-Language-Action Model
- Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success
- Task Robustness via Re-Labelling Vision-Action Robot Data
- VLAs are Confined yet Capable of Generalizing to Novel Instructions
- LangForce: Bayesian Decomposition of Vision Language Action Models via Latent Action Queries
- Flow Matching for Generative Modeling
- Learning to Generalize Across Long-Horizon Tasks from Human Demonstrations
- Unleashing More Actions via Action Compositional Training for VLA Models
- FAST: Efficient Action Tokenization for Vision-Language-Action Models
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving