Security of World-Model-Based Embodied AI: A Lifecycle of Threats, Defenses, and Evaluation
Fazhong Liu, Zhuoyan Chen, Haozhen Tan, Yan Meng, Guoxing Chen, Haojin Zhu
cs.CR, cs.AI
Submitted: 2026-07-30
License: http://creativecommons.org/licenses/by/4.0/
The gist: World models give embodied AI a predictive core: they compress observations into states, simulate action-conditioned futures, and enable planning beyond reactive control.
Terminology
Abstract
World models give embodied AI a predictive core: they compress observations into states, simulate action-conditioned futures, and enable planning beyond reactive control. This predictive layer, however, opens a new security boundary-compromise can propagate from data, sensors, prompts, or feedback into physical action. Rather than treating world models as an isolated component, this survey traces threats across their entire lifecycle-from data construction and representation learning, through state grounding and imagination, to trajectory evaluation, execution, and long-term adaptation via memory and tools. We show that familiar attack families: poisoning, backdoors, adversarial examples, sensor spoofing, prompt injection, trajectory manipulation, and supply-chain attacks take on distinct meanings when they corrupt world states, learned dynamics, affordance estimates, or safety costs. We also highlight a duality: world models can serve as runtime safety shields, yet when compromised or over-trusted they generate predictive safety illusions. The survey offers a lifecycle taxonomy, maps existing attacks to world-model security properties, outlines evaluation protocols for safety failures, and structures defenses across provenance, robust grounding, uncertainty-aware prediction, trajectory gating, feedback auditing, and deployment assurance.
Sources
- World Models
- Mastering Diverse Domains through World Models
- VLM-SAFE: Vision-Language Model-Guided Safety-Aware Reinforcement Learning with World Models for Autonomous Driving
- Genie: Generative Interactive Environments
- RT-1: Robotics Transformer for Real-World Control at Scale
- OpenVLA: An Open-Source Vision-Language-Action Model
- Octo: An Open-Source Generalist Robot Policy
- Towards Robust and Secure Embodied AI: A Survey on Vulnerabilities and Attacks
- What Breaks Embodied AI Security:LLM Vulnerabilities, CPS Flaws,or Something Else?
- Safety of Embodied Navigation: A Survey
- The Safety Challenge of World Models for Embodied AI Agents: A Review
- Safety, Security, and Cognitive Risks in World Models
- Easier Said Than Done: Unpacking Intent-Behavior Gap in Jailbreaking LLM-Based Robots
- When World Models Dream Wrong: Physical-Conditioned Adversarial Attacks against World Models
- State Backdoor: Towards Stealthy Real-world Poisoning Attack on Vision-Language-Action Model in State Space
- ANNIE: Be Careful of Your Robots
- SilentDrift: Exploiting Action Chunking for Stealthy Backdoor Attacks on Vision-Language-Action Models
- TRAP: Tail-aware Ranking Attack for World-Model Planning
- A Framework for Benchmarking and Aligning Task-Planning Safety in LLM-Based Embodied Agents
- Nightshade: Prompt-Specific Poisoning Attacks on Text-to-Image Generative Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs