ContainmentBench: Trace-Based Evaluation of Post-Exposure Containment in Tool-Using LLM Agents
Wenhao Lan, Shan Li, Meiqi Wu, Xinhua Lai, Junbin Yang, Haihua Shen
cs.CR
Submitted: 2026-07-27
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: Tool-using large language model (LLM) agents read untrusted content, maintain memory, delegate tasks, and invoke tools with external side effects.
Terminology
Abstract
Tool-using large language model (LLM) agents read untrusted content, maintain memory, delegate tasks, and invoke tools with external side effects. Terminal attack-success or policy-violation rates do not show what happens between exposure and commit or whether a defense also suppresses authorized actions. We introduce ContainmentBench, a sandboxed benchmark comprising a 504-scenario specification dataset, a shared rollout-trace schema, and stage-scoped metrics for endpoint violations, logged propagation, and explicitly authorized taint-exposed proposals that commit. The main Qwen2.5-7B-Instruct study evaluates seven policy conditions and five seeds, yielding a 17,640-record trace corpus. Across 600 matched active-tainted rollout pairs, no committed policy violation was observed under either taint-only or intent-ledger enforcement. Their execution records nevertheless differed: 441 pairs (73.5%) had different values in a shared 12-field trace summary that includes commit-related diagnostics, and the mean authorized proposal-commit score was 0.164 under taint-only enforcement and 0.857 under intent-ledger enforcement, compared with 0.923 under tool-boundary enforcement. Logged-propagation rankings changed with stage selection and normalization. In a limited set of custom AgentDojo-native workflows, committed violations were observed without defense and were not observed under either evaluated defense. A separate 6,048-rollout Mistral/common-JSON model-interface configuration retained the v1-to-v2 proposal-commit improvement, but committed violations were observed under intent-ledger v2. Equal terminal outcomes do not imply equal containment. The evaluation uses synthetic workflows. The intended intent-ledger mechanism assumes schema-aligned authorization metadata; one public-status task family violates this assumption and is analyzed separately.
Sources
- LLMail-Inject: A Dataset from a Realistic Adaptive Prompt Injection Challenge
- Benchmark Early and Red Team Often: A Framework for Assessing and Managing Dual-Use Hazards of AI Foundation Models
- Ghost in the Agent: Redefining Information Flow Tracking for LLM Agents
- JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
- StruQ: Defending Against Prompt Injection with Structured Queries
- SecAlign: Defending Against Prompt Injection with Preference Optimization
- Securing AI Agents with Information-Flow Control
- Defeating Prompt Injections by Design
- AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
- PIArena: A Platform for Prompt Injection Evaluation
- Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection
- Defending against Adaptive Prompt Injection Attacks via Reasoning-enabled Task Alignment
- AttriGuard: Defeating Indirect Prompt Injection in LLM Agents via Causal Attribution of Tool Invocations
- The Task Shield: Enforcing Task Alignment to Defend Against Indirect Prompt Injection in LLM Agents
- AgentLAB: Benchmarking LLM Agents against Long-Horizon Attacks
- Mitigating Indirect Prompt Injection via Instruction-Following Intent Analysis
- AgentDyn: Are Your Agent Security Defenses Deployable in Real-World Dynamic Environments?
- Hidden in Memory: Sleeper Memory Poisoning in LLM Agents
- Beyond Reproducibility: Towards Security-Aware Evaluation of Research Artifacts
- XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs