Generalizable VLA Finetuning via Representation Anchoring and Language-Action Alignment
summary
The gist
Finetuning vision-language models (VLMs) for robot manipulation via behavior cloning often leads to catastrophic forgetting and language-action misalignment, which this work addresses by proposing
In short
Anchor-Align improves robot manipulation by augmenting behavior cloning with two objectives: Vision-Language Anchoring to preserve pretrained visual representations and Language-Action Alignment to ensure semantic grounding of actions. This method stabilizes finetuning, prevents catastrophic forgetting, and fixes language-action misalignment by linking continuous actions to discrete motion labels.
Key concepts
- Vision-Language Anchoring
- This objective uses a frozen copy of the original Vision-Language Model (VLM) to anchor the representations learned during finetuning. By forcing the backbone's hidden states to align with these frozen states at every layer, it prevents the model from forgetting its general knowledge from pretraining.
- Language-Action Alignment
- This technique converts continuous action targets into discrete labels (like 'up' or 'down') and trains the model to predict these labels based on the robot's observation. This forces a direct connection between what the language predicts and the actual physical motion, correcting semantic errors.
- Catastrophic Forgetting
- This occurs when a neural network, while learning new tasks (finetuning), rapidly loses its ability to perform well on previously learned tasks or general knowledge. Anchor-Align combats this by using anchoring losses to regularize the model's weights toward the original pretrained state.
- Language-Action Misalignment
- This is a problem where the language instructions and the resulting physical actions do not match up correctly. The alignment objective specifically targets this by training a projection layer to map continuous action predictions onto discrete, meaningful motion directions derived from ground-truth data.
Terminology used across episodes
This episode discusses
- Generalizable VLA Finetuning via Representation Anchoring and Language-Action Alignment · Paper Radio
- Qwen2.5-VL Technical Report
- GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
- Evaluating Representational Similarity Measures from the Lens of Functional Correspondence
- RT-1: Robotics Transformer for Real-World Control at Scale
- Towards Synergistic, Generalized, and Efficient Dual-System for Robotic Manipulation
- UniVLA: Learning to Act Anywhere with Task-centric Latent Actions
- Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets
- OpenHelix: A Short Survey, Empirical Analysis, and Open-Source Dual-System VLA Model for Robotic Manipulation
- City Navigation in the Wild: Exploring Emergent Navigation from Web-Scale Knowledge in MLLMs
- Constructive Distortion: Improving MLLMs with Attention-Guided Image Warping · Paper Radio
- Scaling Cross-Embodied Learning: One Policy for Manipulation, Navigation, Locomotion and Aviation
- PaLM-E: An Embodied Multimodal Language Model
- Knowledge Insulating Vision-Language-Action Models: Train Fast, Run Fast, Generalize Better
- LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models
- Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation
- VLA-0: Building State-of-the-Art VLAs with Zero Modification
- Enhancing Generalization in Vision-Language-Action Models by Preserving Pretrained Representations
- Distilling the Knowledge in a Neural Network
- MAPS: Preserving Vision-Language Representations via Module-Wise Proximity Scheduling for Better Vision-Language-Action Generalization
- DreamGen: Unlocking Generalization in Robot Learning through Video World Models
The paper
Generalizable VLA Finetuning via Representation Anchoring and Language-Action Alignment · Read on arXiv
Dwip Dalal, Shivansh Patel, Chahit Jain, Jeonghwan Kim, Utkarsh Mishra, Alex Baratian, Hyeonjeong Ha, Heng Ji
University of Illinois Urbana-Champaign University of Texas at Austin University of California, Irvine
Finetuning a pretrained vision-language model (VLM) on robot demonstrations via behavior cloning (BC) has become the standard recipe for vision-language-action (VLA) policies. However, BC finetuning progressively overwrites the pretrained representations that support visual and semantic generalization. Co-training on web image-text data, a common remedy, applies language and action losses to separate observations, leaving VLAs with language-action misalignment that standard manipulation benchmarks do not expose. We propose Anchor-Align, which augments BC with two objectives: Vision-Language Anchoring distills layer-wise representations from a frozen VLM copy to prevent this drift, while Language-Action Alignment converts each action target into a discrete motion-direction label and jointly trains language and action prediction on the same robot observation. We conduct real-world evaluations across eight manipulation settings on single-arm xArm7 and bimanual YAM robots, using two VLA architectures with regression and flow-matching action heads. Across these settings, Anchor-Align consistently improves over BC on novel targets, layouts, and motion-sensitive bimanual tasks requiring coordinated control. At scale in simulation, we demonstrate consistent improvements on OOD perturbations, perceptual robustness, and long-horizon control across LIBERO-PRO, LIBERO-Plus, and CALVIN, respectively, suggesting that preserving pretrained representations and effective action learning are not fundamentally at odds. Project page: anchoralignvla.github.io
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: Today's paper: "Generalizable VLA Finetuning via Representation Anchoring and Language-Action Alignment".
Dev: Finetuning vision-language models (VLMs) for robot manipulation via behavior cloning often leads to catastrophic forgetting and language-action misalignment, which this work addresses by proposing Anchor-Align,
Rosa: First, who's behind it and why it matters.
Paper summary: Rosa: Well Dev and Taro, we're looking at this paper now titled "Generalizable VLA Finetuning via Representation Anchoring and Language-Action Alignment," and it seems like the core thesis is tackling the problem of catastrophic forgetting that happens when we use standard behavior cloning for vision-language action models.
Dev: Exactly, Rosa; the authors claim that behavior cloning progressively overwrites the pretrained representations that give these VLAs their visual and semantic generalization abilities, which is a real issue when you're fine-tuning on specific robot demonstrations.
Taro: I agree with Dev; standard co-training methods aren't enough because they leave language and action losses separate, leading to language-action misalignment that isn't caught by typical manipulation benchmarks.
Rosa: So, the main idea of Anchor-Align is to fix this by adding two specific objectives on top of the standard behavior cloning loss: Vision-Language Anchoring and Language-Action Alignment.
Dev: Right, that Vision-Language Anchoring part specifically distills layer-wise representations from a frozen VLM copy to actively prevent that representation drift we talked about earlier.
Taro: And what's the second objective, Dev? How does the Language-Action Alignment part help stabilize things when the model gets confused during execution?
Rosa: The Language-Action Alignment converts those continuous action targets into a discrete motion-direction label, like "up" or "down," and then trains the model to predict this label on the same observation it's using for continuous actions.
Dev: That conversion process involves projecting the pre-action hidden state onto a specific vocabulary using a learned projection and the frozen pretrained language head, supervised by that alignment loss.
Taro: That sounds like a smart way to programmatically create targets derived from ground-truth trajectories without needing extra human annotation for every single action.
Rosa: Precisely; this entire Anchor-Align method combines the standard BC loss with these two objectives, and the paper claims it leads to consistent improvements across simulation and real-world experiments.
Dev: The results are pretty compelling, especially when you look at the physical xArm7 robot where they report success rates jumping from twenty-eight percent up to fifty-four percent for one architecture.
Taro: I'm interested in how this affects the system when it encounters situations it hasn't seen before, because they test OOD generalization on benchmarks like LIBERO-PRO and CALVIN.
Rosa: They show improvements not just in simulation, but also under unseen spatial rearrangements, semantic perturbations like picking a pink mug instead of a green one, and even in cluttered scenes during real-world rollouts.
Dev: From an engineering standpoint, that robustness is significant because it means the policy is less brittle when the environment deviates from the training set we gave it.
Taro: So, if we look at the diagnostic value mentioned, they claim this framework gives a direct diagnosis of language-action misalignment in co-trained VLAs and shows better joint alignment scores across functional capabilities.
Rosa: That's a big deal because it moves beyond just seeing that the action is wrong; it tells us *why* the language and action prediction are misaligned, which helps us diagnose the underlying cause.
Dev: Representation analysis confirms this by showing that standard BC causes catastrophic drops in pretrained text representations, but Anchor-Align maintains an average CKA of zero point nine five across layers through that layer-wise distillation.
Taro: It sounds like representation preservation is key here; if you don't keep the underlying knowledge intact, aligning the action will just lead to a misaligned output anyway.
Rosa: And finally, they pointed out that Anchor-Align is more efficient than co-training methods because it requires no external data or annotation and only adds a single inference-only forward pass through a frozen VLM copy.
Dev: That efficiency gain is notable; they say it's about three point four times less overhead per step compared to co-training plus KI, yet the performance is superior.
Taro: Considering how much time we spend gathering and labeling data for those co-training methods, that efficiency in terms of data dependency seems like a major practical advantage.
Rosa: So, to summarize what we've heard about "Generalizable VLA Finetuning via Representation Anchoring and Language-Action Alignment," the paper proposes a method that uses Vision-Language Anchoring to preserve pretrained knowledge and Language-Action Alignment to fix misalignment, resulting in better performance across various challenging real-world scenarios.
Dev: It really seems like they've created a way to stabilize the fine-tuning process without needing massive amounts of new labeled data for every adaptation.
Taro: The implication is that we might be able to adapt these powerful pretrained VLMs much more reliably for complex, real-world robotic tasks than we could with current methods.
Rosa: And I'm thinking about how this translates to the actual deployment in a lab setting; how long do you think this stabilized model can operate reliably outside of the controlled simulation environment?
Dev: That depends on the physical hardware stability and sensor noise, but based on these real-world results, it suggests a much longer operational window before significant performance degradation occurs compared to standard BC.
Taro: If we look at the long-horizon control benchmarks they tested, like CALVIN, that implies this method could be applicable to more complex tasks that require sustained planning rather than just short sequences.
Rosa: It really shows that by focusing on representation stability and semantic grounding through alignment, we can get these VLAs to perform better in the messy reality of physical manipulation.
Conclusion: Rosa: So we're wrapping up our discussion on "Generalizable VLA Finetuning via Representation Anchoring and Language-Action Alignment," which tackles how to make vision-language models work reliably for robots by anchoring their knowledge base and aligning their language outputs with physical actions.
Dev: I'm thinking about the title, Rosa; it sounds like they’re focusing on generalizability, which is crucial when you move from a controlled lab setting to a real environment where things get messy.
Taro: I agree with Dev; that generalizability is what makes these systems useful for true autonomy, especially when the world throws unexpected problems at them.
Rosa: Exactly, and the authors are really focused on how they use these two specific techniques—anchoring and alignment—to solve the problem of forgetting what they learned during pretraining.
Dev: From an engineering standpoint, I'm interested in how this stabilizes the loop rate; if we're adding these extra objectives, does it introduce any noticeable latency or processing overhead that could cause failure modes?
Taro: That’s a valid concern, Dev; the paper does touch on efficiency improvements over co-training methods to keep things lean.
Rosa: And the implication is that this approach could mean we can fine-tune these powerful VLMs much more reliably for complex, real-world tasks without needing tons of new labeled data.
Dev: If that holds true, it changes the deployment timeline significantly because we wouldn't have to spend as much time on manual data collection for every new application.
Taro: It opens up possibilities for more robust autonomy because the system won't just rely on what it memorized, but rather on a more grounded understanding of how language maps to physical movement.
Rosa: So, we’ve seen how the technical details address forgetting and misalignment; now we need to consider the real-world impact of this stabilization method.
Dev: I wonder what kind of long-term operational window we can realistically expect before sensor noise or unexpected environmental changes start pushing these models past their reliable performance threshold.
Taro: That’s a tough question, but if they manage to maintain representation preservation through those layer-wise anchors, it suggests a much longer operational lifespan for these robotic systems.
More episodes
- 2610.12231-Residual Modeling Closes the Regression and Generative Policy Gap in Robot Learning
- 2610.12245-Fixed-Reference Pose Residuals for Measuring Cross-Dataset Cue Transfer in Human-Robot Interaction Anticipation
- 2610.12249-Real-Time Motion Planning with Dynamic Hazards: Classical vs. Learning-Based Methods
- 2610.12272-Walking on Roofs: Exploring the Potential of Walking Robots for Construction Work on Roofs
- 2610.12276-Toward Lunar Legged Robots: Field Deployment Lessons at LUNA
- 2610.12285-PLaW-VLA: Predictive Latent World Modeling for Vision-Language-Action Policies
- 2610.12368-LiteNWM: Efficient Latent World Models for Onboard Visual Navigation in the Wild
- 2610.12435-VioLA: Learning Generalist Humanoid Control Policies from Human Data
- 2610.12404-A Physics-Informed Collision Learning Framework for Collaborative Robot Motion Generation
- 2610.12411-GLIO2: A GPU-Parallelized Tightly-Coupled LiDAR-Inertial-GNSS System for Robust and Real-Time Global Localization and Mapping