ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D

arXiv:2607.19321 · cs.AI, cs.CR, cs.LG · Submitted 2026-07-21 · Read on arXiv

Lena Libon, Ben Rank, Jehyeok Yeon, David Schmotz, Jeremy Qin, Daniel Donnelly, Derck Prinzhorn, Maksym Andriushchenko

cs.AI, cs.CR, cs.LG

Submitted: 2026-07-21

Comments: 50 pages, 12 figures

Code: https://github.com/LLM-QC/judgezoo

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Terminology

Sources

Related papers