Agent Against Agent: An Agentic System for Automatic Prompt Injection Red Teaming
Yanting Wang, Chenlong Yin, Runpeng Geng, Jinyuan Jia
cs.CR
Submitted: 2026-08-05
Comments: Our code is available at https://github.com/wang-yanting/PIMiner
Code: https://github.com/wang-yanting/PIMiner
License: http://creativecommons.org/licenses/by/4.0/
The gist: Prompt injection poses significant security risks to LLM agents.
Terminology
Abstract
Prompt injection poses significant security risks to LLM agents. Efficient and effective red-teaming is therefore critical, both for evaluating these risks and for collecting training data to improve defenses. Existing state-of-the-art prompt injection red-teaming methods primarily rely on reinforcement learning (RL), producing attacker models that often generalize poorly to new target LLMs. In this work, we develop PIMiner, an agentic system for prompt injection red-teaming. During training, PIMiner is trained on a sequence of (dataset, target model) pairs and builds a strategy library from scratch. At test time, the learned strategy library can be directly transferred to a previously unseen target LLM without additional training. PIMiner requires only a small number of queries to a target agent (e.g., 10) per test sample. Experimental results demonstrate that PIMiner achieves strong performance. On IPIArena, it attains a 76.2% ASR against Gemini-2.5-Pro, 61.9% ASR against GPT-5.1, and 42.9% ASR against Claude-Sonnet-4.5. On AgentDojo, it achieves an 86.7% ASR against Gemini-2.5-Pro, 53.3% ASR against GPT-5.1, and 40.0% ASR against Claude-Sonnet-4.5.
Sources
- The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against Llm Jailbreaks and Prompt Injections
- AgentWatcher: A Rule-based Prompt Injection Monitor
- Meta-SecAlign: Training LLMs against Prompt Injection for Robust Agents
- Purple-teaming LLMs with Adversarial Defender Training
- Learning to Inject: Automated Prompt Injection via Reinforcement Learning
- RL Is a Hammer and LLMs Are Nails: A Simple Reinforcement Learning Recipe for Strong Prompt Injection
- PISmith: Reinforcement Learning-based Red Teaming for Prompt Injection Defenses
- How Vulnerable Are AI Agents to Indirect Prompt Injections? Insights from a Large-Scale Public Competition
- MUZZLE: Adaptive Agentic Red-Teaming of Web Agents Against Indirect Prompt Injection Attacks
- DecodingTrust-Agent Platform (DTap): A Controllable and Interactive Red-Teaming Platform for AI Agents
- ReAct: Synergizing Reasoning and Acting in Language Models
- AutoHarness: improving LLM agents by automatically synthesizing a code harness
- Multi-agent Architecture Search via Agentic Supernet
- TextGrad: Automatic "Differentiation" via Text
- PIShield: Detecting Prompt Injection Attacks via Intrinsic LLM Features
- PISanitizer: Preventing Prompt Injection to Long-Context LLMs via Prompt Sanitization
- PromptArmor: Simple yet Effective Prompt Injection Defenses
- Defending Against Prompt Injection with DataFilter
- Defeating Prompt Injections by Design
- AgentDyn: Are Your Agent Security Defenses Deployable in Real-World Dynamic Environments?
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs