Know Your Agent: Reconnaissance-Driven Pentesting of AI Agents
cs.AI, cs.CR, cs.LG
Submitted: 2026-07-22
Updated: 2026-09-15
License: http://creativecommons.org/licenses/by/4.0/
The gist: Traditional pentesting uses reconnaissance at each step to uncover unseen weaknesses, build stronger attacks, and advance the objective; we argue that AI agents require the same treatment.
Terminology
Abstract
Traditional pentesting uses reconnaissance at each step to uncover unseen weaknesses, build stronger attacks, and advance the objective; we argue that AI agents require the same treatment. We formalize agent reconnaissance by modeling the process and identifying the knowledge assets it seeks to extract: what they are, how they are used, and which agent weaknesses they exploit to give adversaries leverage in indirect prompt injection attacks. We instantiate these insights in Know Your Agent (KYA), a framework that automates black-box, reconnaissance-driven pentesting by probing agents, building target profiles, and using those profiles to craft stronger attacks. We evaluate KYA on agent-security benchmarks and a real-world coding agent, and release KYA, its benchmarks, and baseline implementations for reproducibility.
Sources
- ReAct: Synergizing Reasoning and Acting in Language Models
- Prompt Injection attack against LLM-integrated Applications
- InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents
- ChatInject: Abusing Chat Templates for Prompt Injection in LLM Agents
- Prompt Injection Attack to Tool Selection in LLM Agents
- ObliInjection: Order-Oblivious Prompt Injection Attack to LLM Agents with Multi-source Data
- AgentVigil: Generic Black-Box Red-teaming for Indirect Prompt Injection against LLM Agents
- ToolTweak: An Attack on Tool Selection in LLM-based Agents
- Ignore Previous Prompt: Attack Techniques For Language Models
- VeriGrey: Greybox Agent Validation
- SkillJect: Effectively Automating Skill-Based Prompt Injection for Skill-Enabled Agents
- AdapTools: Adaptive Tool-based Indirect Prompt Injection Attacks on Agentic LLMs
- "Real Attackers Don't Compute Gradients": Bridging the Gap Between Adversarial ML Research and Practice
- Universal and Context-Independent Triggers for Precise Control of LLM Outputs
- Trojan's Whisper: Stealthy Manipulation of OpenClaw through Injected Bootstrapped Guidance
- Defending Against Indirect Prompt Injection Attacks With Spotlighting
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection