WMAttack: Automated Attack Search for Adversarial Evaluation of World-Model Agents
cs.LG
Submitted: 2026-05-22
Updated: 2026-09-05
License: http://creativecommons.org/licenses/by/4.0/
The gist: Despite the growing use of world models as decision-making agents, their adversarial robustness remains underexplored due to the lack of dedicated automated evaluation methods.
Terminology
Abstract
Despite the growing use of world models as decision-making agents, their adversarial robustness remains underexplored due to the lack of dedicated automated evaluation methods. A key obstacle is that attack evaluation must be both accurate and efficient: weak manually tuned attacks can overestimate robustness, while exhaustive hyperparameter search is prohibitively expensive because each candidate requires closed-loop rollouts through learned latent dynamics. We introduce WMAttack, an automated attack-search framework for adversarial evaluation of world-model agents. WMAttack formulates robustness evaluation as a finite-budget search over attack configurations, including attack families, perturbation budgets, optimization steps, restarts, and allocation rules. To improve search accuracy, Self-Correcting Attack Search (SCAS) refines the attack proposal distribution using feedback from reward degradation, action instability, runtime cost, and rollout variability. To improve search efficiency, Representation-Guided Attack Retrieval (RGAR) retrieves effective historical configurations from representation-similar tasks, providing a warm start for unseen environments. We provide a theoretical explanation showing that proposal refinement improves finite-budget search when it shifts probability mass toward high-utility attacks. Across Atari and DeepMind Control tasks, WMAttack consistently discovers stronger attacks than the evaluated baselines, improving normalized reward drop from 0.497 to 1.034 on DreamerV3 Atari and from 0.319 to 0.682 on DMC. Ablations further show that RGAR improves initial candidate quality and SCAS improves final attack utility under fixed evaluation budgets.
Sources
- Safe Exploration Using Bayesian World Models and Log-Barrier Optimization
- TRAP: Tail-aware Ranking Attack for World-Model Planning
- The Llama 3 Herd of Models
- When World Models Dream Wrong: Physical-Conditioned Adversarial Attacks against World Models
- Mastering Diverse Domains through World Models
- GAIA-1: A Generative World Model for Autonomous Driving
- Universal Camouflage Attack on Vision-Language Models for Autonomous Driving
- Parallel Rectangle Flip Attack: A Query-based Black-box Attack against Object Detection
- Object Detectors in the Open Environment: Challenges, Solutions, and Outlook
- Improving Adversarial Transferability by Stable Diffusion
- Claudini: Autoresearch Discovers State-of-the-Art Adversarial Attack Algorithms for LLMs
- How Hard is it to Confuse a World Model?
- Uncertainty-aware Latent Safety Filters for Avoiding Out-of-Distribution Failures
- Learning Latent Dynamic Robust Representations for World Models
- DeepMind Control Suite
- Qwen2.5 Technical Report
- WorldBench: Disambiguating Physics for Diagnostic Evaluation of World Models
- Black-Box Adversarial Attack on Vision Language Models for Autonomous Driving
- Text Adversarial Attacks with Dynamic Outputs
- Diversifying the High-level Features for better Adversarial Transferability
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks