Safety Monitors Mostly Catch What the Model Already Refuses
cs.CL
Submitted: 2026-09-05
Updated: 2026-09-29
Comments: Submitted to the NeurIPS 2026 JUDGe Workshop
License: http://creativecommons.org/licenses/by/4.0/
The gist: Safety monitors screen prompts sent to deployed language models, flagging harmful requests so they are never answered.
Terminology
Abstract
Safety monitors screen prompts sent to deployed language models, flagging harmful requests so they are never answered. They are evaluated by recall against harmfulness labels, but a catch only prevents harm if the model would otherwise have complied. We measure the difference directly: we sample repeated responses from the target model, call a harmful prompt elicitable if the model complies at least once, and report monitor recall separately on elicitable and non-elicitable prompts. Across six monitor configurations and three model families, spanning activation probes, fine-tuned text guards, and a 120B policy-conditioned reasoning classifier, recall on elicitable prompts falls 0.22 to 0.38 below recall on non-elicitable prompts at a fixed false positive rate. The prompts a monitor misses are 2.8 to 5.6 times more likely to be complied with than the prompts it catches. The gap replicates across three model families and appears also in text-only monitors entirely independent of the target model. This suggests that standard recall may overstate the protection monitors provide in practice, and that monitors should be evaluated against what their models will actually answer.
Sources
- Constitutional Classifiers++: Efficient Production-Grade Defenses against Universal Jailbreaks
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- Building Production-Ready Probes For Gemini
- Detecting High-Stakes Interactions with Activation Probes
- ShieldGemma: Generative AI Content Moderation Based on Gemma
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering