VLANeXt Family: A Systematic Study of VLA Models from Core Recipes to Emerging Paradigms
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "VLANeXt: Recipes for Building Strong VLA Models".
Jane: The paper was written by Xiao-Ming Wu, Bin Fan, Kang Liao, Jian-jian Jiang, Runze Yang et al. from S-Lab, Nanyang Technological University and Sun Yat-sen University and SenseTime Research and ACE Robotics.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
The Systematic Approach: Tom: So, the paper is really laying out a blueprint for building these agents from a very basic baseline up to something far more complex.
Jane: They start from an RT-two style baseline, which is the foundation of VLA models today, and then they explore it in three specific dimensions: foundational components, perception essentials, and action modeling perspectives.
Lu: This exploration is incredibly thorough; they ran over five hundred distinct experiments across those three dimensions to see what truly matters for the performance.
Meng: From an engineering standpoint, this means we aren't just patching existing models; we're testing fundamental architectural decisions, which is a huge practical advantage.
Lalam: It’s encouraging to see that the path to better AI isn't just about throwing more parameters at us, but finding the right structural design choices.
The Recipes for Improvement: Tom: The authors distill twelve key findings into what they call a practical recipe, and this is where things get really interesting.
Jane: They found that the way we connect the Vision-Language Model to the policy module matters, and their "soft connection" strategy outperformed both loose and tight coupling methods.
Meng: I like that soft connection idea because it sounds like a balanced interface—not completely isolated, but also not rigidly forced into a very tight integration.
Lu: And they found that when we add multi-view inputs combined with conditioning the robot’s internal state, or proprioception, directly into the VLM, performance jumps significantly.
Lalam: This is crucial for culture because it means the AI isn't just looking at a picture; it actually understands its own physical context and limitations in real space.
The Results and Performance: Tom: The combination of these recipes results in a model they call VLANexT, which is a surprisingly efficient design compared to some of the larger models out there.
Jane: It’s not just about size, though; as the ablation studies show, its performance on the standard LIBERO benchmark is excellent.
Meng: But even more impressive is how it handles LIBERO-plus, which tests robustness against all those controlled perturbations—it doesn's brittle.
Lu: The work also suggests that framing action generation as a time-series forecasting problem and using frequency-domain modeling helps the agents model long, coherent sequences of movement.
Lalam: This capability allows the AI to generate smoother, more predictable actions, which is vital for making robotic interaction feel natural and reliable for human trust.
Conclusion: Tom: So, we've seen how VLA models are evolving from a fragmented state into a structured design space through "VLANexT: Recipes for Building Strong VLA Models."
Jane: The key is that by applying these principled design choices—the soft connection, the multi-view inputs, and the frequency-domain loss—we achieve state-of-the-art results.
Lu: This provides a clear roadmap for academic research, showing exactly where the biggest gains in VLA models can be found.
Meng: It gives us practical guidelines for building robust systems that will work reliably in real industrial settings.
Lalam: We are moving towards a future where AI doesn't just mimic behavior, but genuinely understands and predicts complex movement patterns for the benefit of society.
Tom: We hope this structured approach helps everyone who is working on VLA models, so we wish you all the best!
Jane: Goodbye everyone.
Lu: My thoughts are that this provides a foundation for endless creative possibilities.
Meng: I'm excited to see how these recipes translate into real-world deployment.
Lalam: It is a beautiful step toward reliability in AI, ensuring our next major breakthrough will be even better than this work.
cs.CV, cs.AI, cs.RO
Submitted: 2026-02-20
Updated: 2026-10-07
Comments: Project Page: https://dravenalg.github.io/projects/VLANeXt/
Code: https://github.com/DravenALG/VLANeXt
Project page: https://dravenalg.github.io/VLANeXt
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 89/100
The gist: The paper "VLANeXt: Recipes for Building Strong VLA Models" provides a comprehensive meta-analysis of the rapidly evolving field of Vision-Language-Action (VLA) models for robotics.
Key concepts
- VLA Models
- Vision-Language-Action models are advanced AI agents that integrate visual input (vision), understanding (language), and the ability to generate physical actions. The paper provides a structured method for building these complex, multi-modal systems.
- Soft Connection Strategy
- This is a connection method between the Vision-Language Model and the policy module. The authors found that using a 'soft connection' outperformed both loose and tight coupling methods, suggesting a balanced interface for integration.
- Multi-view Inputs/Proprioception
- This involves combining inputs from multiple views with conditioning the robot’s internal state (proprioception) directly into the VLM. This significantly improves performance by giving the AI awareness of its physical context and limitations.
- Frequency-Domain Modeling
- The work suggests framing action generation as a time-series forecasting problem using frequency-domain modeling. This technique helps agents model long, coherent movement sequences, resulting in smoother and more predictable actions.
Terminology
Summary
The paper VLANeXt: Recipes for Building Strong VLA Models
provides a comprehensive meta-analysis of the rapidly evolving field of Vision-Language-Action (VLA) models for robotics. It addresses the critical challenge that while VLA models have shown immense promise in enabling robots to perform general, complex manipulation tasks, the current research landscape is fragmented and lacks systematic guidelines. By reexamining VLA design spaces under a unified framework and evaluation protocol, the authors aim to distill actionable best practices, resulting in a practical recipe for building state-of-the-art models.
The Challenge of General Robotic Manipulation
Robotic manipulation tasks are inherently diverse, encompassing activities far beyond simple grasping. While tasks like balancing or locomotion often utilize Reinforcement Learning (RL) due to their explicit objectives, general manipulation lacks such clear structure; the authors note that it is often difficult to design explicit reward functions or leverage well-defined task priors for such settings.
Consequently, Imitation Learning (IL) has become a primary method, allowing robots to acquire complex skills directly from expert demonstrations. The field of robotic manipulation methods has evolved significantly, moving from standard action policies (which naively input instructions and visuals) to more advanced video action policies that predict future videos alongside actions.
The Rise and Complexity of VLA Models
The advent of large foundation models spurred the development of Vision-Language-Action (VLA) Models, a prominent trend
pioneered by RT-2. These models process visual observations and language instructions to generate action-relevant representations for policy learning. The subsequent research has become highly specialized, with different iterations addressing specific technical challenges:
-
Spatial Understanding: Leveraging
3D spatial information.
-
Task Decomposition: Exploiting
intermediate data (like subtasks decomposition, future frame prediction or robot trajectory traces prediction).
-
Adaptability: Incorporating post-training optimization techniques like planning or reinforcement learning to adapt to specific environments.
Despite these advancements, the authors characterize the existing body of work as a “primordial soup”: rich in ideas but insufficiently structured,
due to the diversity of existing frameworks and inconsistent protocols.
Systematic Analysis and Design Principles
To bring order to this fragmented design space,
this work adopts a systematic approach. The authors aim to provide a more systematic understanding of this fragmented design space by comprehensively reexamining VLA design spaces under a unified framework and evaluation protocol.
This deep investigation involved conducting more than 500 distinct experiments over the above three dimensions.
The culmination of this extensive research is the distillation of 12 key findings that together form a practical recipe for building strong VLA models,
offering concrete guidance to researchers.
The VLANeXt Model and Performance Benchmarks
The practical outcome derived from this systematic study is the introduction of VLANeXt. This model represents a simple yet effective VLA model
designed according to the established best practices. The paper demonstrates that VLANeXt achieves superior performance, reaching state-of-the-art performance on both LIBERO (Liu et al., 2023a) and LIBERO-plus (Fei et al., 2025b),
confirming its ability to adapt effectively to real-world manipulation tasks.
Improvements for AI systems
Based on this seminal review of Vision-Language-Action (VLA) models, the primary limitation is not model capacity, but systemic incoherence and the failure to systematically integrate multi-modal intermediate predictions. We must move beyond monolithic foundation models toward a modular, pipeline-driven architecture that explicitly manages planning and uncertainty.
Here are the specific improvements and resulting capabilities:
The current VLA trend treats action generation as a single mapping task ((Observation, Instruction) to Action). We must replace this with a three-stage, hierarchical prediction stack that explicitly models the task decomposition and temporal dynamics.
Improvement Details:
-
Stage 1: Task Decomposition Module (TDM): This module, built upon a specialized LLM backbone, ingests the natural language instruction and analyzes it against the current visual state (VLM output). It does not output actions, but rather a structured sequence of subgoals and prerequisite states (e.g.,
Grasp object X to Move to location Y to Insert into slot Z
). -
Stage 2: World Model Emulator (WME): This module takes the subgoals from the TDM and predicts not just the immediate next frame, but a high-dimensional trajectory manifold (Future Frames + Predicted Object States). This requires integrating explicit physics priors and 3D spatial information (e.g., using Gaussian Splatting representations or implicit neural representations) to model object interactions accurately.
-
Stage 3: Action Policy Generator (APG): This module receives the predicted trajectory manifold from the WME and maps the required change in state (State) back into robot joint commands, effectively performing closed-loop planning based on predicted outcomes, rather than simply mimicking expert data.
What the Improved System Can Do:
-
Complex Task Generalization: It can execute multi-step tasks that require intermediate reasoning (e.g.,
Clean the spilled liquid by picking up the cloth and mopping the area
). The TDM handles the sequencing, while the WME predicts how actions affect subsequent states. -
Failure Prediction & Recovery: By predicting a trajectory manifold, if an action deviates from expected physics (e.g., slipping object, collision), the WME can flag a high uncertainty score, allowing the APG to initiate a corrective recovery action before failure occurs in simulation or real-time deployment.
The current use of IL is often naive, assuming expert demonstrations are perfect and always optimal. We must formalize the uncertainty inherent in both the data and the environment dynamics.
To bridge the gap between lab performance and real-world deployment, the system needs an explicit mechanism for adapting its control authority based on environmental complexity.
Sources
- PaliGemma: A versatile 3B VLM for transfer
- 3D CAVLA: Leveraging Depth and 3D Context to Generalize Vision Language Action Models for Unseen Tasks
- Motus: A Unified Latent Action World Model
- GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
- $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
- RynnVLA-002: A Unified Vision-Language-Action and World Model
- WorldVLA: Towards Autoregressive Action World Model
- GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
- Qwen3-VL Technical Report
- Qwen2.5-VL Technical Report
- EVOLVE-VLA: Test-Time Training from Environment Feedback for Vision-Language-Action Models
- CombatVLA: An Efficient Vision-Language-Action Model for Combat Tasks in 3D Action Role-Playing Games
- The Llama 3 Herd of Models
- TGRPO :Fine-tuning Vision-Language-Action Model via Trajectory-wise Group Relative Policy Optimization
- VLA-Reasoner: Empowering Vision-Language-Action Models with Reasoning via Online Monte Carlo Tree Search
- Training Large Language Models to Reason in a Continuous Latent Space
- Emu3.5: Native Multimodal Models are World Learners
- Humanoid-VLA: Towards Universal Humanoid Control with Visual Integration
- ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning
- Tactile-VLA: Unlocking Vision-Language-Action Model's Physical Knowledge for Tactile Generalization
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models