APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies

arXiv:2606.12366 · cs.RO · Submitted 2026-06-10 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: Today's paper: "APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies".

Dev: Vision-Language-Action (VLA) models often struggle to generalize to out-of-distribution (OOD) language instructions because continuous action experts, when trained from random initialization on imbalanced data,

Rosa: First, who's behind it and why it matters.

Title and authors: Rosa: Now let's talk about the specifics of this paper, "APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies," and who came up with it. The title itself tells us exactly what they are trying to fix: improving instruction following in VLA policies through action expert pretraining.

Dev: I agree, Rosa; the focus on instruction generalization is key because that’s where most VLA systems fall short when we move from simple lab tasks to real-world scenarios. The authors are Kechun Xu, Zhenjie Zhu, Anzhe Chen, Rong Xiong, and Yue Wang from Zhejiang University and the Zhejiang Humanoid Robot Innovation Center.

Taro: I see the team structure; having researchers from both academia and a dedicated robot innovation center suggests they’re thinking about both the theoretical underpinnings of autonomy and the practical realities of building physical robots.

Rosa: That's right. The authors are clearly trying to bridge that gap, moving beyond just achieving high performance in controlled settings to ensuring the models behave predictably when faced with novel language prompts.

Dev: It’s interesting how they frame it from a Bayesian perspective, which tells us they aren't just throwing random fixes at the data but are trying to build a more principled understanding of how vision, action, and language interact.

Taro: That theoretical framing is important because it helps justify *why* this method works instead of just being an empirical hack that might break later.

Rosa: Exactly. We want to know not just that it works better, but understand the underlying mechanism so we can trust its behavior in the field, which leads us right into what they actually propose doing with this paper.

Dev: So, before we get into the details of how they did it, let's quickly recap: APT is a two-stage training method that uses Bayesian factorization to separate the vision-action prior from the language-conditioned likelihood.

Taro: That separation is where I’m curious; separating those components means you can optimize them independently, which should lead to a more stable overall system when things get complicated.

Rosa: Precisely, Taro; that independence helps prevent one part of the model from corrupting the learning process of another, which is a major hurdle in these coupled systems.

Dev: It sounds like they're essentially trying to build a foundation for motion control first, and then layering the language understanding on top without disturbing that foundation.

The paper's summary: Rosa: We’ve talked about the authors and the title; now let’s get into what they actually summarized in this paper. Essentially, they pinpoint a structural imbalance in VLA data where language is less diverse than visual and action content, which causes continuous action experts to develop visual shortcuts.

Dev: And that leads directly to their core summary: standard training on imbalanced data creates noisy gradients from the action expert that corrupt the VLM backbone because the expert gravitates toward these visual shortcuts instead of learning what the language actually means.

Taro: So, they’re saying that without a specific pretraining strategy, continuous action experts learn to exploit visual cues as a shortcut because those cues are visually rich and abundant in the data.

Rosa: That’s right. They hypothesize that by treating the policy this way—factoring it into a vision-action prior and a language-conditioned likelihood—we can prevent that corruption from happening during the initial learning phase.

Dev: The summary emphasizes that Stage one trains the action expert solely as a VA prior conditioned only on visual tokens from a frozen VLM, which builds this coherent visuomotor manifold without any language influence <ref:2606.12366#pg0>.

Taro: That’s a very specific training recipe; it isolates the motor skills from the linguistic understanding initially, ensuring the physical capabilities are sound before we try to steer them with language.

Rosa: And then Stage two is where they inject those language tokens, training the full VLA policy to align that pre-trained action distribution directly with the desired task instructions <ref:2606.12366#pg0>.

Dev: So, in simple terms, they are first teaching the robot *how* to move based on what it sees, and then teaching it *what* those movements should accomplish based on a language prompt.

Taro: That sounds like a very logical progression for building an autonomous agent; you get the physical skills down first before worrying about the high-level command structure.

The paper's improvements: Rosa: Moving on to what they actually propose as their improvements, APT introduces two key enhancements. First, they suggest using Action Expert Pretraining to create a language-agnostic Vision-Action prior pi p(av) before fine-tuning.

Dev: That pretraining objective is crucial because it’s designed specifically to build that stable visuomotor manifold without the confounding influence of language, which was the main problem in their initial setup.

Taro: And second, they introduce a novel mechanism called "Layer-wise VLM Feature Gated Fusion," which uses trainable scalar gates to modulate how much information from different layers of the VLM backbone gets fed into every self-attention layer.

Rosa: That gating mechanism is smart because it allows the expert to selectively integrate features—it lets it decide whether a feature from a shallow visual layer or a deep semantic layer should influence its action generation at any given moment.

Dev: If that fusion is done correctly, it means the expert can maintain both its own vision-language pathway while still being conditioned by the task language in Stage two which is what they aimed for <ref:2606.12366#pg0>.

Taro: It seems like this dual approach—the prior training and the gated fusion—is what lets them achieve generalization across unseen tasks and compositional instructions, which is a significant capability for autonomy.

Rosa: The results confirm that this combination allows them to handle unseen objects or novel layouts better, showing superior success rates in challenging real-world manipulation scenarios compared to baselines.

Dev: The paper also points out that this method successfully mitigates the issue where language only provides a small amount of additional information beyond what is already visually encoded, by ensuring the prior doesn't condition on language.

Conclusion: Rosa: So, wrapping up our discussion on "APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies," the main implication is that decoupling the visual and action learning via this two-stage Bayesian factorization significantly improves instruction following in VLA models.

Dev: For us as control engineers, the practical implication is that we can trust these policies more when they are deployed because their underlying motion capabilities are more robust against unexpected visual noise during execution.

Taro: From an autonomy standpoint, this means the AI can handle unexpected world behaviors better because it has a stronger grounding in the physical world before applying high-level command logic.

Rosa: It really shows that for continuous action systems, we need to build a solid prior on vision and action independently before trying to teach them complex language tasks.

Dev: I just wonder about the long-term viability; since they noted it doesn't explicitly model long-horizon memory, how does this translate when we consider very complex, multi-step sequences that take hours?

Taro: That’s a fair point; if we need true long-term planning across many hours of interaction, the system will still need an external mechanism for state tracking and memory management.

Rosa: So, to summarize the APT paper, it’s a method that uses pretraining on balanced data and gated fusion to improve instruction generalization in VLA policies. We'll keep an eye on how this concept evolves in future work since they flagged memory as an area for future exploration.

Zhejiang University

cs.RO

Submitted: 2026-06-10

Updated: 2026-10-02

Comments: Accepted by CoRL2026

Code: https://github.com/isaac-sim/IsaacSim

Project page: https://xukechun.github.io/papers/APT

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 93/100

The gist: Vision-Language-Action (VLA) models often struggle to generalize to out-of-distribution (OOD) language instructions because continuous action experts, when trained from random initialization on

Key concepts

Bayesian Factorization
This technique mathematically separates the VLA policy into two independent parts: a Vision-Action (VA) prior and a language-conditioned likelihood. By doing this, the model learns the fundamental physical relationship between seeing an object and performing an action without being distracted by specific language instructions during the initial learning phase.
Action Expert Pretraining (APT)
APT is a two-stage training process. Stage one trains an expert solely on vision and action pairs to build a robust visual-motor understanding. Stage two uses language tokens to steer this pre-trained expert toward specific task goals, ensuring the model can handle new instructions effectively.
Layer-wise VLM Feature Gated Fusion
This is a mechanism in the action expert that allows it to selectively incorporate information from different layers of the Vision Language Model (VLM). A learnable gate decides how much influence each layer's visual and semantic features has on generating an action, allowing the expert to use both shallow spatial details and deep semantic understanding.
Out-of-Distribution (OOD) Generalization
This refers to the model's ability to perform well when given language instructions or tasks it was not specifically trained on. APT improves this by ensuring the action expert learns a general, grounded action prior that is less reliant on specific language cues, leading to better performance on novel instructions.

Terminology

Summary

Vision-Language-Action (VLA) models often struggle to generalize to out-of-distribution (OOD) language instructions because continuous action experts, when trained from random initialization on imbalanced data, develop visual shortcuts that corrupt the Vision Language Model's language capabilities. This paper addresses this by proposing Action Expert Pretraining (APT), a two-stage training method that leverages Bayesian factorization to decouple the vision-action prior from the language-conditioned likelihood, thereby ensuring effective instruction following across unseen tasks and compositional instructions.

The gist

Action expert pretraining on balanced vision-action data improves out-of-distribution language generalization in continuous action VLA policies.

How it works

The core of APT is a Bayesian factorization of the VLA policy into a Vision-Action (VA) prior, πp(av), and a language-conditioned VLA likelihood, L(lv, a). This factorization separates the problem: the VA prior is trained on vision-action pairs alone to build a coherent visuomotor manifold without shortcut incentives, while the VLA likelihood is then fine-tuned with language tokens to steer actions toward specific task instructions.

The training recipe involves two distinct stages:

  1. Stage 1 (VA Prior Pretraining): The action expert is pretrained as a VA prior conditioned solely on visual tokens from a frozen VLM backbone, bypassing the language imbalance by learning πp(av) on balanced vision-action pairs.

  2. Stage 2 (VLA Likelihood Alignment): Language tokens are injected through newly introduced attention layers, forming the likelihood L(lv, a), and the full model is jointly trained to align this pretrained action distribution with task instructions.

Action Expert Design

The proposed action expert is a Transformer-based diffusion model designed with a novel Layer-wise VLM Feature Gated Fusion mechanism. This mechanism injects intermediate features from the Qwen3-VL backbone into every self-attention layer of the action expert using a learnable scalar gate, σ(ˆwi), which modulates the influence of each VLM layer on the action generation process.

The model processes multimodal tokens by concatenating visual tokens (v), language tokens (l), and action tokens (a) into a single sequence and processing them via block-wise causal self-attention. The input to each attention layer is defined as h(i+1)in = h(i)out + σ(ˆwi) · ϕQwen3-VL i(v, l), allowing the expert to assimilate both shallow spatial features and deep semantic features from the VLM while preserving its own vision-language pathway.

Validation and Results

Comprehensive experiments validate that APT achieves consistent gains on unseen instructions and compositional tasks across mainstream VLA architectures, including π and GR00T-style architectures. In simulation benchmarks like LIBERO-PRO, APT consistently outperforms baselines like OpenVLA and π0.5, especially under the Task perturbation where language matters. For rigid object pick-place in real-world settings (Table 3), APT achieves superior success rates across all OOD difficulty levels (UO, UOUC, UOUE) compared to π0.5, demonstrating that action expert pretraining narrows the gap between VLM semantic generalization and grounded action.

Key Findings on Generalization

The analysis confirms that standard VLA training admits a shortcut solution where the conditional mutual information I(a; lv) is upper-bounded by a small constant (H(lv) ≤ ϵ), meaning language provides limited additional information beyond what is visually encoded. APT mitigates this by ensuring the prior πp(av) does not condition on language, and Stage 2 training incentivizes the model to use language for action generation, effectively groundng the prompt. Furthermore, APT shows robust performance in compositional task chaining (e), where it successfully executes sequential sub-tasks without explicit segmentation signals, contrasting sharply with π0.5's failure mode of over-executing the first task while failing to transition.

Architectural Versatility

The effectiveness of APT is shown to generalize across different architectures:

  1. π-style architectures benefit significantly from APT, as the gated fusion design better preserves the VA prior learned in Stage 1 compared to original VLM feature injection.

  2. GR00T-style architectures also show improved generalization when using the two-stage training approach, confirming that action pretraining provides an inductive bias regardless of the specific VLM fusion design.

Limitations

The paper notes that the current design does not explicitly model long-horizon memory, which limits generalization on tasks requiring tracking multi-step progress. Additionally, the evaluation focuses primarily on tabletop manipulation; extending these results to locomotion or mobile manipulation remains an area for future exploration.

Improvements for AI systems

As a fastidious researcher, I have analyzed the provided paper, APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies. The core improvement lies in addressing the structural imbalance in Vision-Language-Action (VLA) data by introducing an explicit pretraining mechanism for continuous action experts.

Here are the specific improvements and capabilities of the improved AI system based on APT:


The proposed system, enhanced with Action Expert Pretraining (APT), moves beyond standard VLA models by decoupling action generation from language grounding through a Bayesian factorization and a two-stage training regimen.

  1. A new pretraining objective is introduced for the continuous action expert using only vision-action pairs (the Vision-Action prior, πp(av)).

  2. The system employs a novel Layer-wise VLM Feature Gated Fusion mechanism, integrating intermediate features from a frozen VLM backbone into every layer of the action expert via trainable gating scalars.

The improved AI system (APT-VLA) can achieve the following specific improvements:

  1. Enhanced generalization to Out-of-Distribution (OOD) language instructions:

  2. Improved robustness against visual shortcuts in continuous manipulation tasks:

  3. Superior handling of compositional and multi-task instructions:

The resulting APT-VLA system will be able to perform the following specific actions and capabilities across various robotic scenarios:

  1. It can reliably follow novel, unseen object references (e.g., picking an eggplant when only trained on grape) even under novel visual layouts and lighting conditions.

  2. It can maintain correct object grounding in cluttered scenes where targets share similar colors or shapes with distractors, avoiding common errors like misgrasping or confusing objects.

  3. It can successfully execute complex, multi-step instructions (e.g., Put the red tea bag into the box followed by Pick up the pepper and place it on the blue plate) without failing to transition between sub-tasks or executing sequential steps in the wrong order (i.e., avoiding failure modes like skipping a required closing action).

  4. It demonstrates superior performance in long-horizon planning tasks, such as those involving push around clutter before a pick-and-place operation, by learning coherent sequences of push, grasp, and place sub-tasks.

Sources

Related papers