BadWAM: When World-Action Models Dream Right but Act Wrong
cs.LG, cs.RO
Submitted: 2026-07-16
Updated: 2026-09-21
License: http://creativecommons.org/licenses/by/4.0/
The gist: World-action models (WAMs) are emerging as a promising foundation for embodied control: rather than predicting actions alone, they learn representations that couple action generation with future
Terminology
Abstract
World-action models (WAMs) are emerging as a promising foundation for embodied control: rather than predicting actions alone, they learn representations that couple action generation with future world prediction. This coupling is often viewed as a source of robustness, interpretability, and safety, as a robot's action can in principle be checked against its imagined future. In this paper, we show that this assumption is fragile. We introduce BadWAM, a unified framework for modeling and evaluating World-Action Drift Attacks: a new class of WAM-specific adversarial attacks that use small visual perturbations to break the alignment between what a WAM imagines and what it executes. BadWAM characterizes this attack surface along two natural criteria: attack strength and stealthiness. When the adversary prioritizes disruption, BadWAM instantiates an action-only adversarial attack, which directly drives the model toward task-failing actions. When the adversary additionally prioritizes stealth, BadWAM instantiates an imagination-preserving adversarial attack, which seeks to induce harmful action shifts while keeping the model's predicted future close to its clean imagination. Together, these two attacks capture a spectrum of WAM-specific failures: from overt action hijacking to stealthier cases where the model appears to imagine a plausible future but executes a desynchronized action. We evaluate BadWAM across different variants of WAMs. Results show that our attacks substantially reduce task success rates under closed-loop execution. For example, our action-only attack reduces the model performance from 96.5% to 43.1% success. The results of our imagination-preserving attack further exposes a WAM-specific vulnerability: moderate future-preserving regularization can maintain strong attack performance while reducing future imagination drift.
Sources
- ABot-M0.5: Unified Mobility-and-Manipulation World Action Model
- RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation
- When World Models Dream Wrong: Physical-Conditioned Adversarial Attacks against World Models
- $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
- Adversarial Patch
- Attacking the Trusted Imagination: Oracle-Level Integrity Attacks on Imagine-then-Act World Models
- ${\pi}_{0.7}$: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities
- Vision-Language-Action Safety: Threats, Challenges, Evaluations, and Mechanisms
- VT-WAM: Visual-Tactile World Action Model for Contact-Rich Manipulation
- Fast-WAM: Do World Action Models Need Test-time Future Imagination?
- JailWAM: Jailbreaking World Action Models in Robot Control
- OA-WAM: Object-Addressable World Action Model for Robust Robot Manipulation
- Understanding and Mitigating the Video-Action Generalization Gap via Temporal Ratio
- BadWorld: Adversarial Attacks on World Models
- World Action Models: A Survey
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks