Your Agentic LLMs Secretly Encode Indirect Prompt-Injection Exposure in Hidden States
Jianshuo Dong, Yiming Liu, Maosen Zhang, Nan Deng, Peng Xu, Xiaoping Zhang, Tianwei Zhang, Jie Zhang, Han Qiu
cs.CR, cs.AI
Submitted: 2026-08-24
Updated: 2026-08-25
Comments: Preprint
Code: https://github.com/jianshuod/IPI-exposure-signal
License: http://creativecommons.org/licenses/by/4.0/
The gist: Agentic LLMs are vulnerable to indirect prompt injection (IPI) attacks, e.g., malicious side-tasks hidden in external tool results.
Terminology
Abstract
Agentic LLMs are vulnerable to indirect prompt injection (IPI) attacks, e.g., malicious side-tasks hidden in external tool results. While many efforts have sought to address the threats, little is known about the internals of agentic LLMs when they are exposed to IPI attacks, a condition which we call IPI exposure. In this paper, we study this problem in depth from three aspects. (1) Probing: Across six models, including the giant 753B-parameter GLM-5.2, simple linear probes trained on pre-generation hidden states can predict LLMs' IPI exposure. These probes achieve 90%+ AUROC on unseen attacks, agent instructions, and task suites; they exhibit high robustness under adaptive attacks and in cross-lingual settings. (2) Defense: Our CoT measurement reveals a recognition--action gap: though models encode such signals, they often fail to translate them into safe actions. We then introduce AGRI, a probe-gated reasoning-based defense that prepends anti-injection reasoning on demand. On difficult AgentDojo settings, AGRI substantially reduces attack success rate, e.g., from 34.6% to 0% on Qwen3.5-27B, while largely maintaining clean-task utility. (3) Explanation: We introduce an analysis framework that identifies natural-language explanations most strongly correlated with probe-captured signals. The resulting profiles differ across models: latent signals can align with either direct IPI-exposure claims or indirect operational cues. Code is available: https://github.com/jianshuod/IPI-exposure-signal.
Sources
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation
- Reasoning Theater: Disentangling Model Beliefs from Chain-of-Thought
- HarmonyGuard: Toward Safety and Utility in Web Agents via Adaptive Policy Enhancement and Dual-Objective Optimization
- LlamaFirewall: An open source guardrail system for building secure AI agents
- Constitutional Classifiers++: Efficient Production-Grade Defenses against Universal Jailbreaks
- Defeating Prompt Injections by Design
- How Vulnerable Are AI Agents to Indirect Prompt Injections? Insights from a Large-Scale Public Competition
- Internal Representations as Indicators of Hallucinations in Agent Tool Selection
- Defending Against Indirect Prompt Injection Attacks With Spotlighting
- Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers
- Measuring Faithfulness in Chain-of-Thought Reasoning
- Commercial LLM Agents Are Already Vulnerable to Simple Yet Dangerous Attacks
- When AUC 0.998 Is Not Enough: A Candidate Evaluation Protocol for Hidden-State Probes of Indirect Prompt Injection in Multimodal Computer-Use Agents
- VIGIL: Defending LLM Agents Against Tool Stream Injection via Verify-Before-Commit
- SafeHarness: Lifecycle-Integrated Security Architecture for LLM-based Agent Deployment
- TraceAegis: Securing LLM-Based Agents via Hierarchical and Behavioral Anomaly Detection
- Prompt Injection attack against LLM-integrated Applications
- The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against Llm Jailbreaks and Prompt Injections
- PromptArmor: Simple yet Effective Prompt Injection Defenses
- The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs