Silent Failures in Agentic Security Evaluation: A Validated Harness for Tool-Call Mediation Under Indirect Prompt Injection
cs.CR, cs.LG
Submitted: 2026-09-26
Updated: 2026-09-26
Terminology
Sources
- The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against Llm Jailbreaks and Prompt Injections
- The Task Shield: Enforcing Task Alignment to Defend Against Indirect Prompt Injection in LLM Agents
- MELON: Provable Defense Against Indirect Prompt Injection Attacks in AI Agents
- IPIGuard: A Novel Tool Dependency Graph-Based Defense Against Indirect Prompt Injection in LLM Agents
- CommandSans: Securing AI Agents with Surgical Precision Prompt Sanitization
- WASP: Benchmarking Web Agent Security Against Prompt Injection Attacks
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs