Safety Monitors Mostly Catch What the Model Already Refuses

arXiv:2609.05797 · cs.CL · Submitted 2026-09-05 · Read on arXiv

cs.CL

Submitted: 2026-09-05

Updated: 2026-09-29

Comments: Submitted to the NeurIPS 2026 JUDGe Workshop

License: http://creativecommons.org/licenses/by/4.0/

The gist: Safety monitors screen prompts sent to deployed language models, flagging harmful requests so they are never answered.

Terminology

Abstract

Safety monitors screen prompts sent to deployed language models, flagging harmful requests so they are never answered. They are evaluated by recall against harmfulness labels, but a catch only prevents harm if the model would otherwise have complied. We measure the difference directly: we sample repeated responses from the target model, call a harmful prompt elicitable if the model complies at least once, and report monitor recall separately on elicitable and non-elicitable prompts. Across six monitor configurations and three model families, spanning activation probes, fine-tuned text guards, and a 120B policy-conditioned reasoning classifier, recall on elicitable prompts falls 0.22 to 0.38 below recall on non-elicitable prompts at a fixed false positive rate. The prompts a monitor misses are 2.8 to 5.6 times more likely to be complied with than the prompts it catches. The gap replicates across three model families and appears also in text-only monitors entirely independent of the target model. This suggests that standard recall may overstate the protection monitors provide in practice, and that monitors should be evaluated against what their models will actually answer.

Sources

Related papers