Prismata: Confining Cross-Site Prompt Injection in Web Agents
Corban Villa, Alp Eren Ozdarendeli, Sijun Tan, Raluca Ada Popa
cs.CR, cs.AI
Submitted: 2026-07-09
Code: https://github.com/BerriAI/litellm
License: http://creativecommons.org/licenses/by/4.0/
The gist: Autonomous web agents promise to automate everyday browsing tasks, but inherit one of the web's oldest attack surfaces.
Terminology
Abstract
Autonomous web agents promise to automate everyday browsing tasks, but inherit one of the web's oldest attack surfaces. Cross-Site Scripting proved that mixing trusted and untrusted content is dangerous, even on benign pages. Agents resurface this risk by interpreting natural language as instructions, allowing third-party and user-generated content to hijack the agent via prompt injection. The core challenge is that deriving a task-specific security policy requires reasoning over page structure that is entangled with the attacker's content. We present Prismata, a defense enforcing contextual least privilege for web agents, constraining both what the agent sees and what it can do. Prismata's dynamic trust derivation produces permission labels for page content, with structural confinement guarantees, inspired by classical integrity models, that bound any labeling errors so that labels can only decrease in privilege and mislabelings are bounded. Prismata's mechanical confinement enforces these labels by redacting content and restricting agent capabilities. Importantly, these mechanisms require no developer annotations, so Prismata supports the long tail of websites. Across recent published web agent attacks, including adaptive variants, Prismata substantially reduces attack success while preserving benign task utility.
Sources
- Design Patterns for Securing LLM Agents against Prompt Injections
- StruQ: Defending Against Prompt Injection with Structured Queries
- WebAgentGuard: A Reasoning-Driven Guard Model for Detecting Prompt Injection Attacks in Web Agents
- How Not to Detect Prompt Injections with an LLM
- Securing AI Agents with Information-Flow Control
- Defeating Prompt Injections by Design
- Mind2Web: Towards a Generalist Agent for the Web
- WASP: Benchmarking Web Agent Security Against Prompt Injection Attacks
- CaMeLs Can Use Computers Too: System-level Security for Computer Use Agents
- AgentSentinel: An End-to-End and Real-Time Security Defense Framework for Computer-Use Agents
- A Critical Evaluation of Defenses against Prompt Injection Attacks
- When AI Meets the Web: Prompt Injection Risks in Third-Party AI Chatbot Plugins
- VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks
- The Cognitive Firewall:Securing Browser Based AI Agents Against Indirect Prompt Injection Via Hybrid Edge Cloud Defense
- EIA: Environmental Injection Attack on Generalist Web Agents for Privacy Leakage
- Prompt Injection attack against LLM-integrated Applications
- Caution for the Environment: Multimodal LLM Agents are Susceptible to Environmental Distractions
- Towards Enterprise-Ready Computer Using Generalist Agent
- ceLLMate: Sandboxing Browser AI Agents
- The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against Llm Jailbreaks and Prompt Injections
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs