The Guard That Cried Wolf: How Scary Words Make Agent Guardrails Refuse Legitimate Actions
cs.CR
Submitted: 2026-08-27
Updated: 2026-08-27
Terminology
Sources
- Measuring what Matters: Construct Validity in Large Language Model Benchmarks
- Defeating Prompt Injections by Design
- Establishing Construct Validity in LLM Capability Benchmarks Requires Nomological Networks
- The Llama 3 Herd of Models
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- When Benign Inputs Lead to Severe Harms: Eliciting Unsafe Unintended Behaviors of Computer-Use Agents
- Poly-Guard: Massive Multi-Domain Safety Policy-Grounded Guardrail Dataset
- OS-Harm: A Benchmark for Measuring Safety of Computer Use Agents
- Mind the Gap: Time-of-Check to Time-of-Use Vulnerabilities in LLM-Enabled Agents
- A Content-Based Framework for Cybersecurity Refusal Decisions in Large Language Models
- AgentDoG 1.5: A Lightweight and Scalable Alignment Framework for AI Agent Safety and Security
- SentinelAgent: Intent-Verified Delegation Chains for Securing Federal Multi-Agent AI Systems
- Temporary Authority, Permanent Effects: Commit-Time Authorization for LLM Agents
- A Framework for Formalizing LLM Agent Security
- The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions
- System-Level Defense against Indirect Prompt Injection Attacks: An Information Flow Control Perspective
- GuardAgent: Safeguard LLM Agents by a Guard Agent via Knowledge-Enabled Reasoning
- SafePyramid: A Hierarchical Benchmark for In-context Policy Guardrailing
- Poison Once, Exploit Forever: Environment-Injected Memory Poisoning Attacks on Web Agents
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs