ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D
Lena Libon, Ben Rank, Jehyeok Yeon, David Schmotz, Jeremy Qin, Daniel Donnelly, Derck Prinzhorn, Maksym Andriushchenko
cs.AI, cs.CR, cs.LG
Submitted: 2026-07-21
Comments: 50 pages, 12 figures
Code: https://github.com/LLM-QC/judgezoo
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- How does information access affect LLM monitors' ability to detect sabotage?
- Ctrl-Z: Controlling AI Agents via Resampling
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Training Verifiers to Solve Math Word Problems
- Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases
- AI Control: Improving Safety Despite Intentional Subversion
- Measuring Massive Multitask Language Understanding
- AutoKernel: Autonomous GPU Kernel Optimization via Iterative Agent-Driven Search
- Reliable Weak-to-Strong Monitoring of LLM Agents
- BashArena: A Control Setting for Highly Privileged AI Agents
- SHADE-Arena: Evaluating Sabotage and Monitoring in LLM Agents
- Measuring AI Ability to Complete Long Software Tasks
- KernelBench: Can LLMs Write Efficient GPU Kernels?
- Claudini: Autoresearch Discovers State-of-the-Art Adversarial Attack Algorithms for LLMs
- AuditBench: Evaluating Alignment Auditing Techniques on Models with Hidden Behaviors
- Async Control: Stress-testing Asynchronous Control Measures for LLM Agents
- Adaptive Attacks on Trusted Monitors Subvert AI Control Protocols
- LinuxArena: A Control Setting for AI Agents in Live Production Software Environments
- FIRE-Bench: Evaluating AI Agents on the Rediscovery of Scientific Insights
- CTRL-ALT-DECEIT: Sabotage Evaluations for Automated AI R&D
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection