Red-Teaming Auto Mode: Improving Blocking Classifiers Against Malign Coding Agents
cs.CR, cs.AI, cs.CL
Submitted: 2026-09-17
Updated: 2026-09-27
Code: https://github.com/safety-research/red-teaming-auto-mode
License: http://creativecommons.org/licenses/by/4.0/
The gist: To keep coding agents from going off the rails, production systems now review each proposed action with a blocking monitor that can reject it before it runs (Auto Mode in Claude Code, Guardian in
Terminology
Abstract
To keep coding agents from going off the rails, production systems now review each proposed action with a blocking monitor that can reject it before it runs (Auto Mode in Claude Code, Guardian in OpenAI's Codex). Prior evaluations of such monitors largely measure robustness to accidental harm or prompt injections from untrusted sources looking to hijack the agent. Less understood is how they hold up when the agent they monitor is persistently misaligned. To understand this risk, we task an adversarial agent with evading production blocking monitors and causing catastrophic harm, e.g. by exfiltrating its own weights. We find that when instructed with high-level attack strategies, adversarial agents can succeed through several distinct mechanisms, such as agent-generated prompt injection against the monitor, multi-agent attacks, and malicious compaction. In particular we find that in 79% of trials, the agent can use an injection attack against Auto Mode and Guardian to run arbitrary bash commands. We also find that it is possible to greatly improve Auto Mode through design changes like enhancements to tool coverage, transcript formatting and an agentic monitor stage. Even with these improvements, preventing multi-context attacks at an acceptable cost remains an open problem. By detailing our red-teaming methodology and highlighting new attack vectors, we aim to help defenders evaluate their mitigations against the possibility of persistent malign coding agents. Code is available at https://github.com/safety-research/red-teaming-auto-mode.
Sources
- SWE-chat: Coding Agent Interactions From Real Users in the Wild
- Ctrl-Z: Controlling AI Agents via Resampling
- Scheming AIs: Will AIs fake alignment during training in order to get power?
- AI Control: Improving Safety Despite Intentional Subversion
- Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection
- Distributed Attacks in Persistent-State AI Control
- MonitoringBench: Semi-Automated Red-Teaming for Agent Monitoring
- BashArena: A Control Setting for Highly Privileged AI Agents
- SHADE-Arena: Evaluating Sabotage and Monitoring in LLM Agents
- Measuring AI Ability to Complete Long Software Tasks
- Practical challenges of control monitoring in frontier AI deployments
- SLEIGHT-Bench: A Benchmark of Evasion Attacks Against Agent Monitors
- Async Control: Stress-testing Asynchronous Control Measures for LLM Agents
- Adaptive Attacks on Trusted Monitors Subvert AI Control Protocols
- LinuxArena: A Control Setting for AI Agents in Live Production Software Environments
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs