Guardrails as Scapegoats: Auditing Unfaithful Safety Refusals in Tool-Augmented LLM Agents

arXiv:2607.19449 · cs.LG, cs.AI, cs.CR · Submitted 2026-07-21 · Read on arXiv

Aarushi Singh

cs.LG, cs.AI, cs.CR

Submitted: 2026-07-21

Comments: 10 pages, 3 figures. Accepted at the ACM KDD 2026 Workshop on Evaluation and Trustworthiness of Agentic AI

Code: https://github.com/langchain-ai/langchain

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Terminology

Sources

Related papers