Will the User Ever Know? Covert Indirect Prompt Injection Attacks on Tool-Using LLM Agents
cs.AI, cs.CL
Submitted: 2026-08-31
Updated: 2026-10-02
Project page: https://yslmoment.github.io/ICoA
Terminology
Sources
- VPI-Bench: Visual Prompt Injection Attacks for Computer-Use Agents
- Mind the GAP: Text Safety Does Not Transfer to Tool-Call Safety in LLM Agents
- ChatInject: Abusing Chat Templates for Prompt Injection in LLM Agents
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- WASP: Benchmarking Web Agent Security Against Prompt Injection Attacks
- The Llama 3 Herd of Models
- Defending Against Indirect Prompt Injection Attacks With Spotlighting
- Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
- GPT-4o System Card
- AgentLAB: Benchmarking LLM Agents against Long-Horizon Attacks
- Refusal-Trained LLMs Are Easily Jailbroken As Browser Agents
- WebGPT: Browser-assisted question-answering with human feedback
- Ignore Previous Prompt: Attack Techniques For Language Models
- Large Language Models can Strategically Deceive their Users when Put Under Pressure
- MUZZLE: Adaptive Agentic Red-Teaming of Web Agents Against Indirect Prompt Injection Attacks
- The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions
- Defending against Indirect Prompt Injection by Instruction Detection
- Qwen3 Technical Report
- ReAct: Synergizing Reasoning and Acting in Language Models
- Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection