SPA: Securing Persistent LLM Agents Across Queries with Plan-First Information-Flow Control
cs.CR, cs.SE
Submitted: 2026-08-27
Updated: 2026-09-04
Comments: Submitted to USENIX Security 2027. 24 pages, 17 figures
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large language model (LLM) agents increasingly operate over untrusted webpages, documents, tools, and persistent states while exercising authority over security-sensitive resources.
Terminology
Abstract
Large language model (LLM) agents increasingly operate over untrusted webpages, documents, tools, and persistent states while exercising authority over security-sensitive resources. Existing defenses typically protect either planning or individual tool interactions, but persistent agents face a broader threat: attacker-controlled data can alter control flow, enter security-sensitive tool arguments, or compromise later queries. We present SPA, a plan-first architecture that secures planning, execution, and cross-query state reuse. SPA invokes the planner once per query to generate a complete executable plan in a declarative domain-specific language, then applies dual-lattice information-flow control to track confidentiality and integrity across explicit data flows and control dependencies. To support persistence without re-exposing untrusted payloads to the planner, SPA stores execution results as labeled artifacts and reveals only semantic metadata during later planning. We evaluate SPA on AgentDojo and AgentDojo-MQ, which is our multi-query extension for measuring secure state reuse and delayed attacks. Under the 'tool knowledge' attack, SPA with information-flow control reduces attack success to zero on AgentDojo and 0.2% on AgentDojo-MQ. Our results show that plan-first execution combined with label-preserving persistence can substantially strengthen persistent LLM agents, while revealing an important security-utility tradeoff introduced by strict integrity enforcement.
Sources
- Securing AI Agents with Information-Flow Control
- Defeating Prompt Injections by Design
- AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
- Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection
- Prompt Flow Integrity to Prevent Privilege Escalation in LLM Agents
- Sophia: A Persistent Agent Framework of Artificial Life
- Les Dissonances: Cross-Tool Harvesting and Polluting in Pool-of-Tools Empowered LLM Agents
- System-Level Defense against Indirect Prompt Injection Attacks: An Information Flow Control Perspective
- The Evolution of Tool Use in LLM Agents: From Single-Tool Call to Multi-Tool Orchestration
- Prompt Injection attack against LLM-integrated Applications
- BrowserAgent: Building Web Agents with Human-Inspired Web Browsing Actions
- Formalizing and Benchmarking Prompt Injection Attacks and Defenses
- InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents
- Prompt Injection Attack to Tool Selection in LLM Agents
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs