Breaking Planner Integrity Boundary: Enviroment State-Text Injection Attack on LLM-Driven Embodied Agents
cs.RO, cs.AI
Submitted: 2026-08-17
Updated: 2026-09-08
Comments: Embodied Agents
Project page: https://code-as-policies.github.io
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large language model (LLM)-driven embodied agents rely on environment states to interpret scenes, generate high-level plans, and drive physical execution, making planner-visible state representations
Terminology
Abstract
Large language model (LLM)-driven embodied agents rely on environment states to interpret scenes, generate high-level plans, and drive physical execution, making planner-visible state representations a critical security boundary. Existing attacks primarily manipulate user instructions, prompt contexts, model behavior, or perceptual inputs, while paying limited attention to whether environment-state text itself can serve as deceptive task evidence and propagate beyond planning to affect execution outcomes. Because embodied tasks are constrained by entity grounding, action preconditions, spatial relations, and environmental constraints, planning deviation alone does not guarantee adversarial execution. To address this gap, we investigate environment-state text as an independent attack surface and present the first closed-loop Environment State-Text Injection (ESTI) attack for LLM-driven embodied agents. Without modifying the original user instruction, model parameters, or executor, ESTI reformulates an adversarial objective as false state evidence compatible with the current environment and influences planning and execution through object properties, spatial relations, affordances, task-stage rules, and execution feedback. We further develop ESTI-Bench to evaluate attack propagation across the planning-to-execution closed loop and compare ESTI with Vanilla IPI, EIRAD, and BADROBOT across ProgPrompt/VirtualHome, VoxPoser/RLBench, and AI2-THOR/iTHOR. ESTI consistently outperforms existing baselines, improving planning-level and execution-level attack success rates by up to 89.32% and 43.69%, respectively. Further analysis shows that grounding, consistency, and executability jointly determine whether manipulated state evidence can propagate through the embodied closed loop and produce verifiable environmental changes.
Sources
- Evaluating Large Language Models Trained on Code
- Do As I Can, Not As I Say: Grounding Language in Robotic Affordances
- Flamingo: a Visual Language Model for Few-Shot Learning
- AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
- Mind2Web: Towards a Generalist Agent for the Web
- PaLM-E: An Embodied Multimodal Language Model
- GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
- Explaining and Harnessing Adversarial Examples
- Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection
- Adversarial Patch
- Language Models are Few-Shot Learners
- Visual Language Maps for Robot Navigation
- CHAI: Command Hijacking against embodied AI
- Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied Agents
- VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models
- Inner Monologue: Embodied Reasoning through Planning with Language Models
- Prompt Injection attack against LLM-integrated Applications
- VIMA: General Robot Manipulation with Multimodal Prompts
- WebGPT: Browser-assisted question-answering with human feedback
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving