When can we trust untrusted monitoring? A safety case sketch across collusion strategies
cs.AI
Submitted: 2026-02-24
Updated: 2026-09-21
Comments: 66 pages, 14 figures, Preprint
Code: https://github.com/UKGovernmentBEIS/inspect_ai
Project page: https://situational-awareness-dataset.org
License: http://creativecommons.org/licenses/by/4.0/
The gist: AIs are increasingly being deployed with greater autonomy and capabilities, which increases the risk that a misaligned AI may be able to cause catastrophic harm.
Terminology
Abstract
AIs are increasingly being deployed with greater autonomy and capabilities, which increases the risk that a misaligned AI may be able to cause catastrophic harm. Untrusted monitoring -- using one untrusted model to oversee another -- is one approach to reducing risk. Justifying the safety of an untrusted monitoring deployment is challenging because developers cannot safely deploy a misaligned model to test their protocol directly. In this paper, we develop upon existing methods for rigorously demonstrating safety based on pre-deployment testing. We relax assumptions that previous AI control research made about the collusion strategies a misaligned AI might use to subvert untrusted monitoring. We develop a taxonomy covering passive self-recognition, causal collusion (hiding pre-shared signals), acausal collusion (hiding signals via Schelling points), and combined strategies. We create a safety case sketch to clearly present our argument, explicitly state our assumptions, and highlight unsolved challenges. We identify conditions under which passive self-recognition could be a more effective collusion strategy than those studied previously. Our work builds towards more robust evaluations of untrusted monitoring.
Sources
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation
- Ctrl-Z: Controlling AI Agents via Resampling
- Looking Inward: Language Models Can Learn About Themselves by Introspection
- Safety Cases: How to Justify the Safety of Advanced AI Systems
- AI Control: Improving Safety Despite Intentional Subversion
- Games for AI Control: Models of Safety Evaluations of AI Deployment Protocols
- Measuring Coding Challenge Competence With APPS
- Authorship Attribution in the Era of LLMs: Problems, Methodologies, and Challenges
- Subversion via Focal Points: Investigating Collusion in LLM Monitoring
- A sketch of an AI control safety case
- How to evaluate control measures for LLM agents? A trajectory from today to superintelligence
- Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs
- Subversion Strategy Eval: Can language models statelessly strategize to subvert control protocols?
- LLM Evaluators Recognize and Favor Their Own Generations
- Idiosyncrasies in Large Language Models
- Adaptive Attacks on Trusted Monitors Subvert AI Control Protocols
- BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection