AcrossWAM1.0:A Modular Latent World-Action Stack for Compact Robot Policies
cs.AI
Submitted: 2026-08-30
Updated: 2026-08-30
License: http://creativecommons.org/licenses/by/4.0/
The gist: Latent world-action models avoid rendering future pixels by predicting an action-relevant visual subgoal in feature space.
Terminology
Abstract
Latent world-action models avoid rendering future pixels by predicting an action-relevant visual subgoal in feature space. LaWAM established this formulation, but its original presentation left the world model, multimodal backbone, and deployment checkpoint tightly coupled. We introduce AcrossWAM1.0, a modularization and scaling study of this latent world-action stack. Rather than presenting latent subgoals as a new algorithm, we make the module boundary explicit: a policy adapter produces latent-action and action-generation contexts; a retained latent world decoder grounds the predicted transition in the current scene;and a flow-matching expert generates continuous action chunks. We further separate training-only teachers from the inference graph and provide a verifiable deployment export. On 2,000 paired LIBERO episodes, replacing a Qwen3-VL-2B backbone with Qwen3.5-0.8B yields 97.45% success versus 98.00% for the 2B model (a-0.55percentage-point difference; exact McNemarp=0.266). This does not prove equivalence, but it meets a prespecified two-point retention criterion. The compact, inference-reachable checkpoint contains 1,472.6M unique parameters, 42.4% fewer than the original 2B policy, while all retained tensors are bitwise identical to the source checkpoint. Cross-family execution is additionally checked with a MiniCPM-V adapter smoke test; closed-loop cross-family transfer remains an open evaluation. AcrossWAM1.0 therefore contributes an auditable software and evaluation boundary for compact latent world-action policies, distinct from LaWAM's original latent-subgoal contribution.
Sources
- Qwen3-VL Technical Report
- Motus: A Unified Latent Action World Model
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
- WorldVLA: Towards Autoregressive Action World Model
- LaWAM: Latent World Action Models for Efficient Dynamics-Aware Robot Policies
- RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation
- VLA-0: Building State-of-the-Art VLAs with Zero Modification
- PriorVLA: Prior-Preserving Adaptation for Vision-Language-Action Models
- Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning
- Causal World Modeling for Robot Control
- Faster-WAM: Do World Action Models Need Deep Action Modules?
- Open X-Embodiment: Robotic Learning Datasets and RT-X Models
- SelfWAM: A Self-Grounded Unified World Action Model for Fast Robot Control
- $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
- DINOv3
- VLA-JEPA: Enhancing Vision-Language-Action Model with Latent World Model
- Fast-WAM: Do World Action Models Need Test-time Future Imagination?
- Faster-WAM: Efficient Inference-Time Future Conditioning for Robust World Action Models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection