ActGuard: Pre-execution Action Auditing against Indirect Prompt Injection in LLM Agents
cs.CR, cs.AI
Submitted: 2026-09-14
Updated: 2026-09-14
Code: https://github.com/binzhwang/ActGuard
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large language model (LLM) agents interact with external environments through tool invocation, but tool outputs can also expose them to indirect prompt injection (IPI) attacks.
Terminology
Abstract
Large language model (LLM) agents interact with external environments through tool invocation, but tool outputs can also expose them to indirect prompt injection (IPI) attacks. Existing defenses mainly rely on prompt hardening, content filtering, pre-generated plans, or permission constraints. These approaches often struggle with complex tasks or over-sanitize external content, making it difficult to balance security and utility. The key challenge is therefore to preserve execution flexibility while precisely identifying and removing the malicious content that actually induces unsafe actions. To address this challenge, we propose ActGuard, a pre-execution action auditing framework. Rather than judging whether external content is inherently suspicious, ActGuard assesses whether it causes the current action to deviate from a locally reasonable expectation. At each step, ActGuard predicts the tools likely to be used by the upcoming action and constructs a local tool prior without constraining the execution trajectory. Before execution, it compares the candidate action against this prior and performs tool-level contrastive analysis and parameter-level evidence localization to identify deviations in tool selection and action parameters. A verifier then examines the localized evidence, masks only spans confirmed as malicious, and regenerates the action from the sanitized context. This design preserves legitimate planning flexibility while minimizing information loss from indiscriminate filtering. We evaluate ActGuard on challenging benchmarks for tool-using agents. Results show that ActGuard reduces attack success rates to a level comparable to state-of-the-art defenses while maintaining task utility close to the no-attack setting, achieving a favorable security-utility trade-off. Our code is publicly available at: https://github.com/binzhwang/ActGuard.
Sources
- LLMail-Inject: A Dataset from a Realistic Adaptive Prompt Injection Challenge
- VPI-Bench: Visual Prompt Injection Attacks for Computer-Use Agents
- Overcoming the Retrieval Barrier: Indirect Prompt Injection in the Wild for LLM Systems
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- From Untrusted Input to Trusted Memory: A Systematic Study of Memory Poisoning Attacks in LLM Agents
- Defeating Prompt Injections by Design
- How Vulnerable Are AI Agents to Indirect Prompt Injections? Insights from a Large-Scale Public Competition
- The Llama 3 Herd of Models
- AttriGuard: Defeating Indirect Prompt Injection in LLM Agents via Causal Attribution of Tool Invocations
- Defending Against Indirect Prompt Injection Attacks With Spotlighting
- CausalArmor: Efficient Indirect Prompt Injection Guardrails via Causal Attribution
- AgentDyn: Are Your Agent Security Defenses Deployable in Real-World Dynamic Environments?
- ReasAlign: Reasoning Enhanced Safety Alignment against Prompt Injection Attack
- RumorSphere: A Framework for Million-scale Agent-based Dynamic Simulation of Rumor Propagation
- Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
- PromptArmor: Simple yet Effective Prompt Injection Defenses
- ICON: Indirect Prompt Injection Defense for Agents based on Inference-Time Correction
- FATH: Authentication-based Test-time Defense against Indirect Prompt Injection Attacks
- AgentWatcher: A Rule-based Prompt Injection Monitor
- AgentVigil: Generic Black-Box Red-teaming for Indirect Prompt Injection against LLM Agents
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs