The Safeguard Worked. Is the LLM System Safer?
cs.CR, cs.AI
Submitted: 2026-09-01
Updated: 2026-09-01
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Safeguards in deployed LLM services are evaluated by refusal, attack success, and policy violation rates.
Terminology
Abstract
Safeguards in deployed LLM services are evaluated by refusal, attack success, and policy violation rates. Those rates characterize how a control performed on the requests it was tested on. A deployment has to answer a different question: how much help with harmful tasks the service still gives an attacker who keeps adapting or finds another way in. We determine what each reported result implies for that question, allowing results from different safeguard families to be compared under one deployment criterion. The evidence requirements are strongly asymmetric. One attack that obtains harmful help from the deployed service suffices to establish that such help remains, and such attacks appear repeatedly in the coded record. Establishing that little remains cannot follow from the safeguard's own numbers alone; it also requires evidence about what the surrounding system still allows after the safeguard performs its local function. Such evidence is supported or derived in only a small minority of the depth-coded claims, and one such claim bounds its scoped residual. A better local score is therefore not, by itself, a stronger claim about the deployment. Safeguard research cannot stop at raising local scores; a gain has to be judged by whether it makes a deployed system any safer.
Sources
- TeleAI-Safety: A comprehensive LLM jailbreaking benchmark towards attacks, defenses, and evaluations
- A Systematic Literature Review on LLM Defenses Against Prompt Injection and Jailbreaking: Expanding NIST Taxonomy
- Securing AI Agents with Information-Flow Control
- Reward Model Ensembles Help Mitigate Overoptimization
- Adversarial Paraphrasing: A Universal Attack for Humanizing AI-Generated Text
- Mechanistic Unlearning: Robust Knowledge Unlearning and Editing via Mechanistic Localization
- Bias Runs Deep: Implicit Reasoning Biases in Persona-Assigned LLMs
- SoK: Systematizing LLM Prompt Security: Taxonomies, Datasets, and Unified Evaluation of Attacks and Defenses
- GuidedBench: Measuring and Mitigating the Evaluation Discrepancies of In-the-wild LLM Jailbreak Methods
- TrustLLM: Trustworthiness in Large Language Models
- SceneJailEval: A Scenario-Adaptive Multi-Dimensional Framework for Jailbreak Evaluation
- Inference-Time Intervention: Eliciting Truthful Answers from a Language Model
- LLM Cyber Evaluations Don't Capture Real-World Risk
- Secret Collusion among AI Agents: Multi-Agent Deception via Steganography
- Detecting High-Stakes Interactions with Activation Probes
- Red Teaming the Mind of the Machine: A Systematic Evaluation of Prompt Injection and Jailbreak Vulnerabilities in LLMs
- Identifying the Risks of LM Agents with an LM-Emulated Sandbox
- Beyond I'm Sorry, I Can't: Dissecting Large Language Model Refusal
- Progent: Securing AI Agents with Privilege Control
- PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs