Same Scene, Different Task: Skill Alignment for Compositional Generalization in VLAs
summary
The gist
Fine-tuned Vision-Language-Action (VLA) models often struggle to generalize to skill combinations absent from their fine-tuning demonstrations, even when every constituent skill has been demonstrated.
In short
CRAFT is a counterfactual supervision method that helps Vision-Language-Action (VLA) models perform skill combinations they haven't seen during training. It uses existing demonstrations of individual skills to train the model on new, unseen sequences by creating 'counterfactual pairs.' This allows VLAs to generalize better to complex tasks without needing new data for every possible combination.
Key concepts
- Compositional Generalization
- This is the challenge where a VLA model knows how to perform individual skills but fails when those skills are put together in a new order or with different objects. The paper addresses this by training models to handle these combinations they haven't explicitly seen before.
- Counterfactual Pairs
- These are pairs created by keeping the visual observation the same but changing the instruction to request an undemonstrated skill combination. This setup helps train the model on what action to take when given a novel instruction, even though no direct action target was shown for that specific pair.
- Skill Representation ($z_{skill}$)
- This is a learned feature that captures the essence of a specific skill, independent of the current visual state. It is trained to be reusable across different executions of the same skill, allowing the model to apply knowledge about one skill to a new context effectively.
- Skill Query Tokens ($Q_{skill}$)
- These are learnable tokens added to the VLM sequence that focus on understanding the required skill. They attend to both visual and text information simultaneously to generate a 'skill representation' that guides the action prediction process.
Terminology used across episodes
This episode discusses
- Same Scene, Different Task: Skill Alignment for Compositional Generalization in VLAs · Paper Radio
- Do As I Can, Not As I Say: Grounding Language in Robotic Affordances
- GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
- pi 0: A Vision-Language-Action Flow Model for General Robot Control
- pi 0.5: a Vision-Language-Action Model with Open-World Generalization
- ReSteer: Quantifying and Refining the Steerability of Multitask Robot Policies
- When Vision Overrides Language: Evaluating and Mitigating Counterfactual Failures in VLAs
- Efficient Data Collection for Robotic Manipulation via Compositional Generalization
- CAST: Counterfactual Labels Improve Instruction Following in Vision-Language-Action Models
- Enhancing Generalization in Vision-Language-Action Models by Preserving Pretrained Representations
- PriorVLA: Prior-Preserving Adaptation for Vision-Language-Action Models
- LoRA: Low-Rank Adaptation of Large Language Models
- OpenVLA: An Open-Source Vision-Language-Action Model
- Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success
- Task Robustness via Re-Labelling Vision-Action Robot Data
- VLAs are Confined yet Capable of Generalizing to Novel Instructions
- LangForce: Bayesian Decomposition of Vision Language Action Models via Latent Action Queries
- Flow Matching for Generative Modeling
- Learning to Generalize Across Long-Horizon Tasks from Human Demonstrations
- Unleashing More Actions via Action Compositional Training for VLA Models
- FAST: Efficient Action Tokenization for Vision-Language-Action Models
The paper
Same Scene, Different Task: Skill Alignment for Compositional Generalization in VLAs · Read on arXiv
Taegeun Yang, Youngju Na, Yoonki Cho, Sung-Eui Yoon
Korea Advanced Institute of Science and Technology (KAIST
Vision-language-action (VLA) models often struggle to generalize to skill combinations absent from their fine-tuning demonstrations, even when every constituent skill has been demonstrated. We focus on a vision shortcut as one failure mode: during fine-tuning, visual observations can serve as a proxy for the instruction, so a policy may execute a demonstrated combination associated with similar observations rather than the instructed combination. This motivates training with counterfactual pairs formed by holding a demonstration observation fixed while changing the instruction to specify an undemonstrated combination. These pairs, however, lack corresponding demonstrated action targets. Crucially, the currently required skill has already been demonstrated, but actions from those executions cannot serve as direct targets because the same skill can require different actions across observations. We propose CRAFT, which transfers supervision from demonstrated executions of the required skill to counterfactual pairs using skill representations that can be reused across executions of the same skill. Across three VLA models and two simulation benchmarks, CRAFT improves success on undemonstrated combinations while maintaining high success on demonstrated ones; it also improves compositional generalization on a real robot. Project website: https://taegeunyang.github.io/craft/
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: Today's paper: "Same Scene, Different Task".
Dev: Fine-tuned Vision-Language-Action (VLA) models often struggle to generalize to skill combinations absent from their fine-tuning demonstrations, even when every constituent skill has been demonstrated.
Rosa: First, who's behind it and why it matters.
Title and authors: Rosa: So we're starting with "Same Scene, Different Task: Skill Alignment for Compositional Generalization in VLAs," and I'm really interested in what this paper is tackling because compositional generalization is such a big hurdle for robots right now.
Dev: Yeah, it sounds like they are diving into that exact problem where models get stuck on combinations they haven't explicitly seen during training. What’s the core idea behind their approach to solving that?
Taro: Basically, the paper points out that when a task is made of smaller skills, like picking and placing, and you only show them those skills separately in demonstrations, the AI struggles when you ask it to do a new combination like picking one object and placing it on a different surface.
Rosa: Exactly. The title suggests they are trying to fix this by aligning skill representations so the model doesn't rely too heavily on just looking at what's in front of it instead of paying attention to the specific instruction for that moment.
Dev: I saw their summary mentions that they use existing demonstrations of constituent skills to train VLAs to execute skill combinations absent from the original demonstrations without having to collect new data. That sounds like a smart way around the data collection bottleneck we usually face in these VLA systems.
Taro: That’s interesting because it avoids the huge cost of generating entirely new trajectories for every possible combination, which is what some other approaches have tried to do by creating synthetic demonstrations.
Rosa: Right, and their methodology seems to hinge on using skill representations that can be reused across different executions of the same basic skill, while also allowing action predictions to change based on the observation. It sounds like they are trying to separate *what* skill is being done from *where* exactly it's happening in the scene.
Dev: And I noticed they introduce these learnable query tokens, Qskill and Qstate, where the skill queries look at both image and text tokens to build a skill representation, while state queries only look at the image for a state representation. That sounds like a clever way to give the model different ways to interpret visual information.
Taro: I think that separation is key; if we can isolate the pure skill concept from the specific visual context, it might help when things get messy in real-world scenarios where unexpected stuff happens.
Title and authors: Rosa: And then they layer on three training objectives: a standard flow-matching loss to supervise the action expert with what they already have, an Lskill objective to make those representations reusable across different skill executions, and an Lcf for counterfactual skill-representation alignment.
Dev: That counterfactual supervision part sounds crucial; it’s not just about matching demonstrations anymore; it’s about training the representation to correctly infer the required skill even when the instruction is changed in a way that's absent from the original data.
Taro: If that Lcf objective works as they suggest, it means we can give them supervision for a new combination using an old demonstration of just one of its parts, which is exactly what we need for robustness.
Rosa: So, to sum up what we've heard about "Same Scene, Different Task: Skill Alignment for Compositional Generalization in VLAs," they are proposing a method that uses existing skill demonstrations to teach the VLA model how to handle combinations it hasn't seen before by focusing on learning reusable skill representations.
Dev: It seems like the practical application is making these models significantly better at generalizing, even when the instruction requires chaining skills in a novel way. I wonder if this level of generalization translates well when we look outside the clean simulation environment.
Taro: That’s my main concern about deployment; if it works perfectly in simulation but fails when the world misbehaves, that’s where we have to be careful. The paper doesn't explicitly detail how this handles novel physical interactions outside of their structured benchmarks.
Rosa: That brings us right to the next point: the improvements they suggest for this framework, which seem designed to make it even more robust and versatile than what was presented in the core paper.
Dev: The improvements focus on making those skill representations truly reusable across different executions of that same skill, while also ensuring they are distinct enough for different skills. It’s about managing that trade-off between reusability and specialization.
Taro: And I like the idea of using counterfactual supervision to train the representation to reflect a required skill under a changed instruction, rather than just matching the input observation. That should give it more reasoning power when faced with ambiguity.
Title and authors: Rosa: It seems they are building on their initial concept by adding mechanisms that enforce this skill identity separation and use the counterfactual pairs actively during training, which should make the system more flexible in real-world deployment situations where instructions might be phrased differently.
Dev: From an engineering standpoint, I’m focused on how this affects latency and loop rate; if these new representations add too much complexity to the inference pipeline, we could see performance hit our required real-time constraints. I'd want to see details on the computational overhead of those Qskill and Qstate queries.
Taro: That's a fair point, Dev; any extra layers in the policy architecture need to be efficient enough so they don't introduce unacceptable delays in a dynamic environment where speed matters for safety.
Rosa: So, we’ve covered the core idea of how "Same Scene, Different Task: Skill Alignment for Compositional Generalization in VLAs" uses counterfactual supervision to tackle skill combination failures by learning reusable skill representations. We're now looking at how the authors suggest pushing this framework further with specific architectural improvements.
Dev: And those improvements seem aimed squarely at solving the generalization gap while keeping the system computationally feasible, which is a tough balancing act for any VLA system.
Taro: I think if they can successfully decouple skill identity from specific observations, that opens up possibilities for more adaptable autonomy when things go wrong in unexpected ways.
Rosa: Indeed, it sounds like the paper lays a solid foundation by showing that we don't need new data for every combination if we use these clever alignment techniques to train the underlying skill understanding.
Dev: So, as we wrap up this discussion on "Same Scene, Different Task: Skill Alignment for Compositional Generalization in VLAs," it seems like they offer a concrete way to improve how models handle complex task sequencing without needing massive amounts of new demonstration data.
Taro: I think the real impact here is showing that existing demonstrations can be leveraged far more effectively than we currently do when dealing with sequential actions.
Rosa: Absolutely, and this work provides valuable insight into structuring the training supervision to encourage better compositional understanding in VLA models. We’ll take a quick pause before we look at what these findings actually mean for the wider field of robotics.
The paper's summary: Rosa: So, we've been hearing about "Same Scene, Different Task: Skill Alignment for Compositional Generalization in VLAs," and now I want to talk about what that summary actually means for us as field roboticists and engineers.
Dev: It seems the core of the paper boils down to using existing demonstrations of individual skills to train a VLA model so it can successfully execute entirely new skill combinations it hasn't seen before, all without needing new data collection.
Taro: That’s what interests me about the counterfactual supervision part; essentially, they're using what they already know about one skill to teach the model how to handle another when the instruction changes in a way that wasn't in the original training set.
Rosa: Exactly, it tackles that vision shortcut problem where models get confused by visual cues instead of following the actual sequence of operations specified in a complex task.
Dev: From an engineering standpoint, I’m thinking about how this impacts our loop rates; if these learned skill representations are too slow or complex to process, we won't get the real-time performance we need for deployment.
Taro: That's a valid concern, Dev; the paper suggests they introduce mechanisms to make those representations reusable across different executions of the same skill, which should theoretically keep the computational load manageable while boosting generalization.
Rosa: It sounds like their results show that this approach improves success on combinations we haven't demonstrated—like picking an object and placing it in a new spot—while still keeping high success on the combinations they actually showed us during fine-tuning.
Dev: The evaluation results, especially when you look at real-robot testing, are pretty impressive; they show significant jumps in performance on those undemonstrated tasks compared to standard fine-tuning methods.
Taro: I'm really excited about the implication here for autonomy; if we can reliably generalize from known skills, it means we can build systems that are much more flexible and less brittle when faced with unexpected scenarios outside of the lab.
Rosa: That’s what makes me keen to know how long this works in practice; does this generalization hold up when the environment gets messy or when things move unexpectedly?
Dev: The paper’s limitations are clear, though; they state that this formulation assumes a fixed sequence of operations and that every single skill needed for an undemonstrated combination was present in the original demonstrations.
Taro: That limitation is important; so if we have a completely novel operation or a brand new skill entirely, this specific framework won't cover it without modification.
Rosa: Well, that means we need to watch how the authors propose expanding this concept to handle those more open-ended tasks and varying instruction sequences in the future.
The paper's improvements: Rosa: So, we've discussed how "Same Scene, Different Task: Skill Alignment for Compositional Generalization in VLAs" uses counterfactual supervision to teach models about skill combinations without needing new data.
Dev: Right, and now we're looking at the suggested improvements they propose to make this approach even more robust and effective for real-world applications.
Taro: I'm looking at the points where they suggest learning a reusable skill representation that is distinct from others, which should really help when the model has to switch between different skills quickly.
Rosa: It sounds like they're focusing on how to make that skill identity separation clearer so the VLA can properly condition its actions based on what it needs to do at that specific moment.
Dev: I’m focused on the computational cost here; these additional learning objectives, like Lskill and Lcf, introduce more complexity into the training process, so I want to know how much overhead we’re looking at during inference when running these improved models.
Taro: The counterfactual supervision objective they mention is really interesting because it trains the skill representation to reflect the exact skill required by a changed instruction under fixed reference states.
Rosa: That sounds like it gives us more reasoning power; it means when an instruction is phrased in a way we haven't seen, the AI has a better chance of inferring the correct underlying action requirement.
Dev: If this improved system can successfully maintain high success on demonstrated combinations while substantially increasing success on undemonstrated ones, that’s a significant step toward deploying these agents in more complex environments.
Taro: I think this directly addresses the brittle nature of current VLA systems; instead of failing completely when an instruction shifts slightly, it should be able to adapt its behavior based on its learned skill representations.
Rosa: That opens up possibilities for much more versatile robotic assistants, capable of handling a wider variety of tasks without needing entirely new training datasets for every single variation.
Dev: We need to see if this adaptability translates into reliable, low-latency execution during actual operation, because a clever representation that takes too long to compute is just useless in a fast-paced physical world.
Taro: The paper flags a limitation where this method assumes the task sequence is fixed and every skill needed was shown in the original demonstrations; so we still need to address how it handles genuinely novel or entirely unrepresented skills.
Rosa: Exactly, so the future work needs to focus on extending this framework beyond fixed sequences and incorporating mechanisms for learning those entirely new skills dynamically.
Conclusion: Rosa: So, to wrap things up, we've seen how "Same Scene, Different Task: Skill Alignment for Compositional Generalization in VLAs" uses counterfactual supervision to train VLA models on skill combinations they haven't explicitly seen before.
Dev: It really shows a way to boost the system’s ability to generalize from learned skills without needing massive amounts of new demonstration data.
Taro: I think the real value is in how it handles uncertainty; by aligning skill representations this way, we get better reasoning when the world throws us an instruction that doesn't perfectly match our original training examples.
Rosa: And that’s exactly where we need to watch: whether this robustness holds up when we take these agents out of the lab and into truly unstructured environments for extended periods.
Dev: From my side, I’m still checking the computational overhead; if the complexity of those skill representations causes any significant latency spikes during high-speed execution, that’s a failure mode we can't ignore.
Taro: That adaptability is what makes this concept important for autonomy; if we can make systems more flexible when things misbehave, it moves us closer to truly reliable navigation and manipulation in the wild.
Rosa: It’s exciting to think about how this could help build robots that are less prone to breaking down when faced with a slightly different task sequence.
Dev: We need more data on the long-term stability of these learned skill representations under continuous operation, not just benchmark success rates in simulation or controlled settings.
Taro: So, we’re left wondering how far this technique can stretch before it hits its limits when dealing with completely new physical interactions that aren't covered by the existing skill demonstrations.
Rosa: That’s the open question for our field: what are the necessary next steps to push this framework into truly general-purpose autonomy?
Dev: I'm curious about how they plan to make these learned skills more resilient against external disturbances that might affect the state representation during operation.
More episodes
- 2610.12154-Stochastic Distribution Network Reconfiguration under Load Uncertainty
- 2607.00148-3D Point World Models: Point Completion Enables More Accurate Dynamics Learning
- 2607.02403-ACID: Action Consistency via Inverse Dynamics for Planning with World Models
- 2510.26623-A Sliding-Window Filter for Online Continuous-Time Continuum Robot State Estimation
- 2406.13267-The Kinetics Observer: A Tightly Coupled Estimator for Legged Robots
- 2511.02147-Census-Based Population Autonomy For Distributed Robotic Teaming
- 2603.08260-Seed2Scale: A Self-Evolving Data Engine with Parallel Worlds Expansion for Scalable Robot Learning
- 2602.14032-RoboAug: One Annotation to Hundreds of Scenes via Region-Contrastive Data Augmentation for Robotic Manipulation
- 2602.15397-ActionCodec: What Makes for Good Action Tokenizers
- 2607.01819-Koopman operator theory: fundamentals, control, and applications