Improving Evaluation Realism with Inference-Time Compute and Deployment Scaffolds
cs.AI, cs.CL, cs.LG
Submitted: 2026-09-02
Updated: 2026-09-02
Comments: 70 pages, 43 figures, 4 tables (13 figures in the main text). Under review at NeurIPS 2026. Code: https://github.com/meridianlabs-ai/petri_dish and https://github.com/AxelAhlqvist1995/petri-bon ; reproduction assets: https://github.com/AxelAhlqvist1995/petri-realism-reproduction
Code: https://github.com/meridianlabs-ai/petri_dish
Project page: https://www-cdn.anthropic.com/8b8380204f74670be75e81c820ca8dda846ab289.pdf
License: http://creativecommons.org/licenses/by/4.0/
The gist: A core obstacle to alignment evaluation is evaluation awareness: capable models can tell when they are being tested rather than deployed, weakening the conclusions a safety evaluation can support.
Terminology
Abstract
A core obstacle to alignment evaluation is evaluation awareness: capable models can tell when they are being tested rather than deployed, weakening the conclusions a safety evaluation can support. We present two techniques that make simulated alignment evaluations harder to distinguish from real deployments. Our first technique, critique refinement, spends additional inference-time compute on each simulator action: the simulator generates multiple candidate actions, refines them using feedback from an instance of the target model on how to make them more realistic, and continues the evaluation with the most deployment-like candidate. Our second technique, DISH (Deployment-Imitating SWE-Agent Harness), wraps the target in an agent harness, reducing the gap between simulated and real deployment environments in coding settings. We test the techniques on multiple target models and find that they compose: applying both yields larger realism gains than either alone. Our results show that automated approaches can improve the realism of alignment evaluations, and that these improvements use additional compute more effectively than making the audits longer.
Sources
- LURE: Live-Usage Replay Evaluations for Reducing Evaluation Awareness
- Evaluating whether AI models would sabotage AI safety research
- Gram: Assessing sabotage propensities via automated alignment auditing
- Self-Refine: Iterative Refinement with Self-Feedback
- Frontier Models are Capable of In-context Scheming
- Large Language Models Often Know When They Are Being Evaluated
- Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
- UK AISI Alignment Evaluation Case-Study
- WildChat: 1M ChatGPT Interaction Logs in the Wild
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection