Generalizable VLA Finetuning via Representation Anchoring and Language-Action Alignment

arXiv:2607.13429 · cs.RO, cs.CV · Submitted 2026-07-15 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: Today's paper: "Generalizable VLA Finetuning via Representation Anchoring and Language-Action Alignment".

Dev: Finetuning vision-language models (VLMs) for robot manipulation via behavior cloning often leads to catastrophic forgetting and language-action misalignment, which this work addresses by proposing Anchor-Align,

Rosa: First, who's behind it and why it matters.

Paper summary: Rosa: Well Dev and Taro, we're looking at this paper now titled "Generalizable VLA Finetuning via Representation Anchoring and Language-Action Alignment," and it seems like the core thesis is tackling the problem of catastrophic forgetting that happens when we use standard behavior cloning for vision-language action models.

Dev: Exactly, Rosa; the authors claim that behavior cloning progressively overwrites the pretrained representations that give these VLAs their visual and semantic generalization abilities, which is a real issue when you're fine-tuning on specific robot demonstrations.

Taro: I agree with Dev; standard co-training methods aren't enough because they leave language and action losses separate, leading to language-action misalignment that isn't caught by typical manipulation benchmarks.

Rosa: So, the main idea of Anchor-Align is to fix this by adding two specific objectives on top of the standard behavior cloning loss: Vision-Language Anchoring and Language-Action Alignment.

Dev: Right, that Vision-Language Anchoring part specifically distills layer-wise representations from a frozen VLM copy to actively prevent that representation drift we talked about earlier.

Taro: And what's the second objective, Dev? How does the Language-Action Alignment part help stabilize things when the model gets confused during execution?

Rosa: The Language-Action Alignment converts those continuous action targets into a discrete motion-direction label, like "up" or "down," and then trains the model to predict this label on the same observation it's using for continuous actions.

Dev: That conversion process involves projecting the pre-action hidden state onto a specific vocabulary using a learned projection and the frozen pretrained language head, supervised by that alignment loss.

Taro: That sounds like a smart way to programmatically create targets derived from ground-truth trajectories without needing extra human annotation for every single action.

Rosa: Precisely; this entire Anchor-Align method combines the standard BC loss with these two objectives, and the paper claims it leads to consistent improvements across simulation and real-world experiments.

Dev: The results are pretty compelling, especially when you look at the physical xArm7 robot where they report success rates jumping from twenty-eight percent up to fifty-four percent for one architecture.

Taro: I'm interested in how this affects the system when it encounters situations it hasn't seen before, because they test OOD generalization on benchmarks like LIBERO-PRO and CALVIN.

Rosa: They show improvements not just in simulation, but also under unseen spatial rearrangements, semantic perturbations like picking a pink mug instead of a green one, and even in cluttered scenes during real-world rollouts.

Dev: From an engineering standpoint, that robustness is significant because it means the policy is less brittle when the environment deviates from the training set we gave it.

Taro: So, if we look at the diagnostic value mentioned, they claim this framework gives a direct diagnosis of language-action misalignment in co-trained VLAs and shows better joint alignment scores across functional capabilities.

Rosa: That's a big deal because it moves beyond just seeing that the action is wrong; it tells us *why* the language and action prediction are misaligned, which helps us diagnose the underlying cause.

Dev: Representation analysis confirms this by showing that standard BC causes catastrophic drops in pretrained text representations, but Anchor-Align maintains an average CKA of zero point nine five across layers through that layer-wise distillation.

Taro: It sounds like representation preservation is key here; if you don't keep the underlying knowledge intact, aligning the action will just lead to a misaligned output anyway.

Rosa: And finally, they pointed out that Anchor-Align is more efficient than co-training methods because it requires no external data or annotation and only adds a single inference-only forward pass through a frozen VLM copy.

Dev: That efficiency gain is notable; they say it's about three point four times less overhead per step compared to co-training plus KI, yet the performance is superior.

Taro: Considering how much time we spend gathering and labeling data for those co-training methods, that efficiency in terms of data dependency seems like a major practical advantage.

Rosa: So, to summarize what we've heard about "Generalizable VLA Finetuning via Representation Anchoring and Language-Action Alignment," the paper proposes a method that uses Vision-Language Anchoring to preserve pretrained knowledge and Language-Action Alignment to fix misalignment, resulting in better performance across various challenging real-world scenarios.

Dev: It really seems like they've created a way to stabilize the fine-tuning process without needing massive amounts of new labeled data for every adaptation.

Taro: The implication is that we might be able to adapt these powerful pretrained VLMs much more reliably for complex, real-world robotic tasks than we could with current methods.

Rosa: And I'm thinking about how this translates to the actual deployment in a lab setting; how long do you think this stabilized model can operate reliably outside of the controlled simulation environment?

Dev: That depends on the physical hardware stability and sensor noise, but based on these real-world results, it suggests a much longer operational window before significant performance degradation occurs compared to standard BC.

Taro: If we look at the long-horizon control benchmarks they tested, like CALVIN, that implies this method could be applicable to more complex tasks that require sustained planning rather than just short sequences.

Rosa: It really shows that by focusing on representation stability and semantic grounding through alignment, we can get these VLAs to perform better in the messy reality of physical manipulation.

Conclusion: Rosa: So we're wrapping up our discussion on "Generalizable VLA Finetuning via Representation Anchoring and Language-Action Alignment," which tackles how to make vision-language models work reliably for robots by anchoring their knowledge base and aligning their language outputs with physical actions.

Dev: I'm thinking about the title, Rosa; it sounds like they’re focusing on generalizability, which is crucial when you move from a controlled lab setting to a real environment where things get messy.

Taro: I agree with Dev; that generalizability is what makes these systems useful for true autonomy, especially when the world throws unexpected problems at them.

Rosa: Exactly, and the authors are really focused on how they use these two specific techniques—anchoring and alignment—to solve the problem of forgetting what they learned during pretraining.

Dev: From an engineering standpoint, I'm interested in how this stabilizes the loop rate; if we're adding these extra objectives, does it introduce any noticeable latency or processing overhead that could cause failure modes?

Taro: That’s a valid concern, Dev; the paper does touch on efficiency improvements over co-training methods to keep things lean.

Rosa: And the implication is that this approach could mean we can fine-tune these powerful VLMs much more reliably for complex, real-world tasks without needing tons of new labeled data.

Dev: If that holds true, it changes the deployment timeline significantly because we wouldn't have to spend as much time on manual data collection for every new application.

Taro: It opens up possibilities for more robust autonomy because the system won't just rely on what it memorized, but rather on a more grounded understanding of how language maps to physical movement.

Rosa: So, we’ve seen how the technical details address forgetting and misalignment; now we need to consider the real-world impact of this stabilization method.

Dev: I wonder what kind of long-term operational window we can realistically expect before sensor noise or unexpected environmental changes start pushing these models past their reliable performance threshold.

Taro: That’s a tough question, but if they manage to maintain representation preservation through those layer-wise anchors, it suggests a much longer operational lifespan for these robotic systems.

Dwip Dalal, Shivansh Patel, Chahit Jain, Jeonghwan Kim, Utkarsh Mishra, Alex Baratian, Hyeonjeong Ha, Heng Ji

University of Illinois Urbana-Champaign University of Texas at Austin University of California, Irvine

cs.RO, cs.CV

Submitted: 2026-07-15

Updated: 2026-09-28

Comments: Code: https://github.com/dwipddalal/Anchor-Align

Code: https://github.com/huggingface/lerobot

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 92/100

The gist: Finetuning vision-language models (VLMs) for robot manipulation via behavior cloning often leads to catastrophic forgetting and language-action misalignment, which this work addresses by proposing

Key concepts

Vision-Language Anchoring
This objective uses a frozen copy of the original Vision-Language Model (VLM) to anchor the representations learned during finetuning. By forcing the backbone's hidden states to align with these frozen states at every layer, it prevents the model from forgetting its general knowledge from pretraining.
Language-Action Alignment
This technique converts continuous action targets into discrete labels (like 'up' or 'down') and trains the model to predict these labels based on the robot's observation. This forces a direct connection between what the language predicts and the actual physical motion, correcting semantic errors.
Catastrophic Forgetting
This occurs when a neural network, while learning new tasks (finetuning), rapidly loses its ability to perform well on previously learned tasks or general knowledge. Anchor-Align combats this by using anchoring losses to regularize the model's weights toward the original pretrained state.
Language-Action Misalignment
This is a problem where the language instructions and the resulting physical actions do not match up correctly. The alignment objective specifically targets this by training a projection layer to map continuous action predictions onto discrete, meaningful motion directions derived from ground-truth data.

Terminology

Summary

Finetuning vision-language models (VLMs) for robot manipulation via behavior cloning often leads to catastrophic forgetting and language-action misalignment, which this work addresses by proposing Anchor-Align, a method that augments standard behavior cloning with two objectives to preserve pretrained representations and ensure semantic grounding in action.

The gist

Anchor-Align augments behavior cloning with two objectives: VisionLanguage Anchoring distills layer-wise representations from a frozen VLM copy to prevent this drift, while Language-Action Alignment converts each action target into a discrete motion-direction label and jointly trains language and action prediction on the same robot observation.

How it works

The method combines the standard behavior cloning loss with two novel objectives: Vision-Language Anchoring and Language-Action Alignment, optimized jointly via the total loss:

Ltotal = Laction + λanchor Lanchor + λalign Lalign (Equation 1).

  1. Vision-Language Anchoring distills representations from a frozen pretrained VLM copy to prevent catastrophic forgetting. This is achieved by anchoring hidden states of the vision and text tokens between the backbone and anchor VLM at every decoder layer, using the per-layer anchoring loss L(u)anchor (Equation 2).

  2. Language-Action Alignment converts each ground-truth demonstration trajectory into a discrete language label and trains the model to predict this label on the same observation it acts on. This involves projecting the pre-action hidden state onto a vocabulary using a learned projection and the frozen pretrained language head, supervised by Equation 5: Lalign = CE(aˆlang, alang).

Key Objectives

The two objectives serve distinct roles in stabilizing and aligning the model:

(Vision-Language Anchoring)

This objective prevents catastrophic forgetting by anchoring the VLA’s backbone VLM with a frozen copy of the same VLM (the anchor VLM), which processes the same input batch in parallel. This distills original pretrained representations into the trainable backbone, where anti-forgetting methods regularize finetuned models toward their pretrained copy in weight space or representation space.

(Language-Action Alignment)

This objective addresses language-action misalignment by programmatically converting continuous action targets into a discrete motion-direction label (e.g., up, down) and training the model to predict this label on the same robot observation where it predicts the continuous actions. The alignment targets are derived from ground-truth action trajectories with no human annotation through steps involving average chunking, filtering, and discretization to generate a discrete direction label alang.

Experimental Validation

Anchor-Align was evaluated across simulation benchmarks (LIBERO-PRO, LIBERO-Plus, CALVIN) and real-world experiments on an xArm7 robot. The results demonstrate consistent improvements:

  1. On the physical xArm7 robot, Anchor-Align improves real-robot success on both architectures: (28% → 54% and 37% → 60%).

  2. In simulation, it improves OOD generalization on LIBERO-PRO and LIBERO-Plus benchmarks, including position swap tests.

  3. It shows consistent improvements across perceptual robustness, long-horizon control (CALVIN), unseen spatial rearrangements, semantic perturbations (e.g., picking a pink mug when trained on a green one), and cluttered scenes in real-world rollouts.

Analysis of Performance

The analysis confirms that both anchoring and alignment are necessary:

(Necessity of Both)

Ablation studies show that both objectives independently improve performance over standard behavior cloning. Combining them in Anchor-Align VLA exceeds either alone, indicating they play complementary roles: anchoring preserves the representational substrate that alignment leverages.

(Diagnostic Value)

The Language-Action Alignment framework provides the first direct diagnosis of language-action misalignment in co-trained VLAs, showing that better alignment improves action accuracy. Furthermore, analysis using an extended diagnostic dataset (covering Motion Direction, Task Completion, Grasp, and Orientation axes) reveals that co-trained VLAs often exhibit poor alignment across all axes; Anchor-Align VLA substantially improves joint alignment scores on these functional capabilities.

(Representation Preservation)

Representation analysis confirms that standard BC catastrophically destroys pretrained text representations (CKA drops to 0.34), whereas Anchor-Align recovers nearperfect preservation through layer-wise distillation, maintaining an average CKA of 0.95 across layers, while simultaneously attaining the highest action decodability (peak R2 = 0.60 at layer 22).

(Efficiency)

Anchor-Align is more efficient than co-training methods. It requires no external data or annotation and adds only a single inference-only forward pass through a frozen copy of the pretrained VLM, resulting in +3.4× less overhead per step compared to co-training + KI, while achieving superior downstream performance.

Improvements for AI systems

Based on the Anchor-Align method and its experimental results, here are specific improvements that can be implemented in Vision-Language-Action (VLA) systems:


  1. Improving Generalization to Novel Visual Configurations (Robustness):

  2. Enhancing Semantic Reasoning Under Instruction Shifts (Instruction Following):

  3. Achieving Higher Fidelity in Long-Horizon and Complex Tasks (Long-Horizon Control):

  4. Enabling Real-World Deployment with Enhanced Efficiency and Speed:

  5. Improving Generalization to Novel Visual Configurations (Robustness):

The system can robustly handle unseen object positions, novel object types, and significant scene layout changes without catastrophic forgetting of previously learned skills. By using the frozen VLM copy as an anchor and distilling its representations layer-wise during finetuning, the VLA retains its broad spatial priors (e.g., understanding affordances like pick up, place on) even when the exact scene configuration is perturbed (position swap, object swap) or when visual attributes change (lighting changes).

  1. Enhancing Semantic Reasoning Under Instruction Shifts (Instruction Following):

The system can reliably follow complex, paraphrased, or context-dependent instructions in real-world environments. The Language-Action Alignment objective forces the language understanding head to predict a discrete motion direction based on the current observation, thereby preventing the language output from contradicting the required physical action. This allows VLAs to succeed when instructions are rephrased (e.g., pick up X becomes grab Y) or when they must ground abstract concepts in a cluttered scene.

  1. Achieving Higher Fidelity in Long-Horizon and Complex Tasks (Long-Horizon Control):

The system can execute multi-step, chained tasks with high accuracy across extended rollouts (e.g., 5+ consecutive instructions). The combination of preserving pretrained representations ensures that small grounding errors do not compound over time. Furthermore, the alignment objective ensures that the model maintains coherent motion plans throughout the entire trajectory, leading to more decisive and efficient action vectors at critical phases (like grasping), resulting in faster and more natural robot behavior.

  1. Enabling Real-World Deployment with Enhanced Efficiency and Speed:

The system can achieve significant speedups in real-world execution time compared to standard finetuned models. By producing higher-magnitude, more decisive action vectors during critical phases (like grasp), the model reduces the number of small, tentative corrective steps required by the low-level controller. This results in a nearly 1.7x faster average rollout time and tighter completion distributions, making the policy suitable for high-throughput applications where cycle time is critical.

Abstract

Finetuning a pretrained vision-language model (VLM) on robot demonstrations via behavior cloning (BC) has become the standard recipe for vision-language-action (VLA) policies. However, BC finetuning progressively overwrites the pretrained representations that support visual and semantic generalization. Co-training on web image-text data, a common remedy, applies language and action losses to separate observations, leaving VLAs with language-action misalignment that standard manipulation benchmarks do not expose. We propose Anchor-Align, which augments BC with two objectives: Vision-Language Anchoring distills layer-wise representations from a frozen VLM copy to prevent this drift, while Language-Action Alignment converts each action target into a discrete motion-direction label and jointly trains language and action prediction on the same robot observation. We conduct real-world evaluations across eight manipulation settings on single-arm xArm7 and bimanual YAM robots, using two VLA architectures with regression and flow-matching action heads. Across these settings, Anchor-Align consistently improves over BC on novel targets, layouts, and motion-sensitive bimanual tasks requiring coordinated control. At scale in simulation, we demonstrate consistent improvements on OOD perturbations, perceptual robustness, and long-horizon control across LIBERO-PRO, LIBERO-Plus, and CALVIN, respectively, suggesting that preserving pretrained representations and effective action learning are not fundamentally at odds. Project page: anchoralignvla.github.io

Sources

Related papers