APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
summary
The gist
Vision-Language-Action (VLA) models often struggle to generalize to out-of-distribution (OOD) language instructions because continuous action experts, when trained from random initialization on
In short
The paper proposes Action Expert Pretraining (APT), a two-stage training method to improve how Vision-Language-Action (VLA) models follow instructions in new situations. It decouples the visual-action relationship from language conditioning by first training an action expert on vision and action data alone, then fine-tuning it with language tokens. This results in better generalization to unseen tasks.
Key concepts
- Bayesian Factorization
- This technique mathematically separates the VLA policy into two independent parts: a Vision-Action (VA) prior and a language-conditioned likelihood. By doing this, the model learns the fundamental physical relationship between seeing an object and performing an action without being distracted by specific language instructions during the initial learning phase.
- Action Expert Pretraining (APT)
- APT is a two-stage training process. Stage one trains an expert solely on vision and action pairs to build a robust visual-motor understanding. Stage two uses language tokens to steer this pre-trained expert toward specific task goals, ensuring the model can handle new instructions effectively.
- Layer-wise VLM Feature Gated Fusion
- This is a mechanism in the action expert that allows it to selectively incorporate information from different layers of the Vision Language Model (VLM). A learnable gate decides how much influence each layer's visual and semantic features has on generating an action, allowing the expert to use both shallow spatial details and deep semantic understanding.
- Out-of-Distribution (OOD) Generalization
- This refers to the model's ability to perform well when given language instructions or tasks it was not specifically trained on. APT improves this by ensuring the action expert learns a general, grounded action prior that is less reliant on specific language cues, leading to better performance on novel instructions.
Terminology used across episodes
This episode discusses
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies · Paper Radio
- pi 0: A Vision-Language-Action Flow Model for General Robot Control
- GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
- LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization
- A Taxonomy for Evaluating Generalist Robot Manipulation Policies
- Seeing to Act, Prompting to Specify: A Bayesian Factorization of Vision Language Action Policy
- When Vision Overrides Language: Evaluating and Mitigating Counterfactual Failures in VLAs
- FAST: Efficient Action Tokenization for Vision-Language-Action Models
- Gemini Robotics: Bringing AI into the Physical World
- pi 0.5: a Vision-Language-Action Model with Open-World Generalization
- A Survey on Vision-Language-Action Models: An Action Tokenization Perspective
- OpenHelix: A Short Survey, Empirical Analysis, and Open-Source Dual-System VLA Model for Robotic Manipulation
- Octo: An Open-Source Generalist Robot Policy
- Open X-Embodiment: Robotic Learning Datasets and RT-X Models
- RoboMIND: Benchmark on Multi-embodiment Intelligence Normative Data for Robot Manipulation
- AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems
- RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation
- RoboVerse: Towards a Unified Platform, Dataset and Benchmark for Scalable and Generalizable Robot Learning
- Qwen3-VL Technical Report
- PaliGemma: A versatile 3B VLM for transfer
- Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success
The paper
APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies · Read on arXiv
Zhejiang University
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: Today's paper: "APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies".
Dev: Vision-Language-Action (VLA) models often struggle to generalize to out-of-distribution (OOD) language instructions because continuous action experts, when trained from random initialization on imbalanced data,
Rosa: First, who's behind it and why it matters.
Title and authors: Rosa: Now let's talk about the specifics of this paper, "APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies," and who came up with it. The title itself tells us exactly what they are trying to fix: improving instruction following in VLA policies through action expert pretraining.
Dev: I agree, Rosa; the focus on instruction generalization is key because that’s where most VLA systems fall short when we move from simple lab tasks to real-world scenarios. The authors are Kechun Xu, Zhenjie Zhu, Anzhe Chen, Rong Xiong, and Yue Wang from Zhejiang University and the Zhejiang Humanoid Robot Innovation Center.
Taro: I see the team structure; having researchers from both academia and a dedicated robot innovation center suggests they’re thinking about both the theoretical underpinnings of autonomy and the practical realities of building physical robots.
Rosa: That's right. The authors are clearly trying to bridge that gap, moving beyond just achieving high performance in controlled settings to ensuring the models behave predictably when faced with novel language prompts.
Dev: It’s interesting how they frame it from a Bayesian perspective, which tells us they aren't just throwing random fixes at the data but are trying to build a more principled understanding of how vision, action, and language interact.
Taro: That theoretical framing is important because it helps justify *why* this method works instead of just being an empirical hack that might break later.
Rosa: Exactly. We want to know not just that it works better, but understand the underlying mechanism so we can trust its behavior in the field, which leads us right into what they actually propose doing with this paper.
Dev: So, before we get into the details of how they did it, let's quickly recap: APT is a two-stage training method that uses Bayesian factorization to separate the vision-action prior from the language-conditioned likelihood.
Taro: That separation is where I’m curious; separating those components means you can optimize them independently, which should lead to a more stable overall system when things get complicated.
Rosa: Precisely, Taro; that independence helps prevent one part of the model from corrupting the learning process of another, which is a major hurdle in these coupled systems.
Dev: It sounds like they're essentially trying to build a foundation for motion control first, and then layering the language understanding on top without disturbing that foundation.
The paper's summary: Rosa: We’ve talked about the authors and the title; now let’s get into what they actually summarized in this paper. Essentially, they pinpoint a structural imbalance in VLA data where language is less diverse than visual and action content, which causes continuous action experts to develop visual shortcuts.
Dev: And that leads directly to their core summary: standard training on imbalanced data creates noisy gradients from the action expert that corrupt the VLM backbone because the expert gravitates toward these visual shortcuts instead of learning what the language actually means.
Taro: So, they’re saying that without a specific pretraining strategy, continuous action experts learn to exploit visual cues as a shortcut because those cues are visually rich and abundant in the data.
Rosa: That’s right. They hypothesize that by treating the policy this way—factoring it into a vision-action prior and a language-conditioned likelihood—we can prevent that corruption from happening during the initial learning phase.
Dev: The summary emphasizes that Stage one trains the action expert solely as a VA prior conditioned only on visual tokens from a frozen VLM, which builds this coherent visuomotor manifold without any language influence <ref:2606.12366#pg0>.
Taro: That’s a very specific training recipe; it isolates the motor skills from the linguistic understanding initially, ensuring the physical capabilities are sound before we try to steer them with language.
Rosa: And then Stage two is where they inject those language tokens, training the full VLA policy to align that pre-trained action distribution directly with the desired task instructions <ref:2606.12366#pg0>.
Dev: So, in simple terms, they are first teaching the robot *how* to move based on what it sees, and then teaching it *what* those movements should accomplish based on a language prompt.
Taro: That sounds like a very logical progression for building an autonomous agent; you get the physical skills down first before worrying about the high-level command structure.
The paper's improvements: Rosa: Moving on to what they actually propose as their improvements, APT introduces two key enhancements. First, they suggest using Action Expert Pretraining to create a language-agnostic Vision-Action prior pi p(av) before fine-tuning.
Dev: That pretraining objective is crucial because it’s designed specifically to build that stable visuomotor manifold without the confounding influence of language, which was the main problem in their initial setup.
Taro: And second, they introduce a novel mechanism called "Layer-wise VLM Feature Gated Fusion," which uses trainable scalar gates to modulate how much information from different layers of the VLM backbone gets fed into every self-attention layer.
Rosa: That gating mechanism is smart because it allows the expert to selectively integrate features—it lets it decide whether a feature from a shallow visual layer or a deep semantic layer should influence its action generation at any given moment.
Dev: If that fusion is done correctly, it means the expert can maintain both its own vision-language pathway while still being conditioned by the task language in Stage two which is what they aimed for <ref:2606.12366#pg0>.
Taro: It seems like this dual approach—the prior training and the gated fusion—is what lets them achieve generalization across unseen tasks and compositional instructions, which is a significant capability for autonomy.
Rosa: The results confirm that this combination allows them to handle unseen objects or novel layouts better, showing superior success rates in challenging real-world manipulation scenarios compared to baselines.
Dev: The paper also points out that this method successfully mitigates the issue where language only provides a small amount of additional information beyond what is already visually encoded, by ensuring the prior doesn't condition on language.
Conclusion: Rosa: So, wrapping up our discussion on "APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies," the main implication is that decoupling the visual and action learning via this two-stage Bayesian factorization significantly improves instruction following in VLA models.
Dev: For us as control engineers, the practical implication is that we can trust these policies more when they are deployed because their underlying motion capabilities are more robust against unexpected visual noise during execution.
Taro: From an autonomy standpoint, this means the AI can handle unexpected world behaviors better because it has a stronger grounding in the physical world before applying high-level command logic.
Rosa: It really shows that for continuous action systems, we need to build a solid prior on vision and action independently before trying to teach them complex language tasks.
Dev: I just wonder about the long-term viability; since they noted it doesn't explicitly model long-horizon memory, how does this translate when we consider very complex, multi-step sequences that take hours?
Taro: That’s a fair point; if we need true long-term planning across many hours of interaction, the system will still need an external mechanism for state tracking and memory management.
Rosa: So, to summarize the APT paper, it’s a method that uses pretraining on balanced data and gated fusion to improve instruction generalization in VLA policies. We'll keep an eye on how this concept evolves in future work since they flagged memory as an area for future exploration.
More episodes
- 2610.12154-Stochastic Distribution Network Reconfiguration under Load Uncertainty
- 2607.00148-3D Point World Models: Point Completion Enables More Accurate Dynamics Learning
- 2607.02403-ACID: Action Consistency via Inverse Dynamics for Planning with World Models
- 2510.26623-A Sliding-Window Filter for Online Continuous-Time Continuum Robot State Estimation
- 2406.13267-The Kinetics Observer: A Tightly Coupled Estimator for Legged Robots
- 2511.02147-Census-Based Population Autonomy For Distributed Robotic Teaming
- 2603.08260-Seed2Scale: A Self-Evolving Data Engine with Parallel Worlds Expansion for Scalable Robot Learning
- 2602.14032-RoboAug: One Annotation to Hundreds of Scenes via Region-Contrastive Data Augmentation for Robotic Manipulation
- 2602.15397-ActionCodec: What Makes for Good Action Tokenizers
- 2607.01819-Koopman operator theory: fundamentals, control, and applications