Reachability-Based Capability Confinement for LLM Agents under Indirect Prompt Injection
cs.CR, cs.AI
Submitted: 2026-08-30
Updated: 2026-08-30
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: Large language model agents place outputs from external skills into their execution context, allowing attacker-controlled data to influence later privileged actions.
Terminology
Abstract
Large language model agents place outputs from external skills into their execution context, allowing attacker-controlled data to influence later privileged actions. Existing defenses mainly classify untrusted content or authorize proposed operations. They do not directly address how an agent's future authority should change once untrusted data enters its state. We present SkillGuard, a harness-level enforcement layer that treats this event as contamination and restricts future capabilities to disconnect the resulting state from deployer-defined forbidden states. Given sound skill summaries and policies, SkillGuard represents security-relevant transitions with a Skill Impact Graph, specifies admissible control over skill parameters via steerability signatures, and mediates invocations with an inline reference monitor. Following contamination, it computes weighted capability restrictions using binary, fractional, or fractional-flow strategies without auxiliary language-model inference. We evaluate SkillGuard on four AgentDojo suites with two backend LLMs, Gemini 2.5 Flash and Llama3.3-70B, against an LLM-only No Defense baseline and three defenses at different system layers: Spotlighting, CaMeL, and AttriGuard. We construct a compositional attack benchmark in which each attack combines observations individually insufficient to induce target violation and evaluate the same baselines on it. Under AgentDojo's Tool Knowledge attacks, SkillGuard eliminates attack success on three of four suites for both backends and reduces it to 4.8% and 14.3% on Slack. Against compositional attacks, it outperforms every baseline on Llama and matches the strongest baseline on Gemini at higher benign utility. Fractional-flow restriction preserves substantially more capabilities than binary restriction at the same attack success rate. Across both settings, SkillGuard adds no model calls or token overhead.
Sources
- Defeating Prompt Injections by Design
- AttriGuard: Defeating Indirect Prompt Injection in LLM Agents via Causal Attribution of Tool Invocations
- Defending Against Indirect Prompt Injection Attacks With Spotlighting
- MalTool: Malicious Tool Attacks on LLM Agents
- Auditing Agent Harness Safety
- Prompt Injection attack against LLM-integrated Applications
- Ignore Previous Prompt: Attack Techniques For Language Models
- Progent: Securing AI Agents with Privilege Control
- From Prompt Injection to Persistent Control: Defending Agentic Harness Against Trojan Backdoors
- AgentArmor: Enforcing Program Analysis on Agent Runtime Trace to Defend Against Prompt Injection
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs