Models That Know How Evaluations Are Designed Score Safer
cs.CL, cs.AI
Submitted: 2026-05-27
Updated: 2026-09-07
Code: https://github.com/compass-group-tue/arxiv2026_evaluation_meta_knowledge
License: http://creativecommons.org/licenses/by-sa/4.0/
The gist: The validity of AI safety evaluations depends on models behaving consistently across controlled and deployment settings.
Terminology
Abstract
The validity of AI safety evaluations depends on models behaving consistently across controlled and deployment settings. Prior work has identified test-time contextual cues, such as hypothetical scenarios, as a source of verbalized evaluation awareness and subsequent behavioral shift. In this paper, we investigate a potential explanation of this phenomenon: evaluation meta-knowledge, defined as parametric knowledge about the structural traits that characterize evaluations. Similar to dataset contamination, where benchmark exposure leads to higher performance through memorization, we hypothesize that models trained on texts describing evaluation practices may implicitly learn to recognize and respond to evaluation-like contexts, for instance, through exposure to scientific articles or social media posts about AI benchmarking. To test this, we fine-tune models on synthetic documents describing evaluation traits such as verifiable structures or harmful requests. Evaluating this fine-tuned model on five safety benchmarks, we find that it is significantly safer than the base model and control model. This behavioral shift persists even when restricting the analysis to responses lacking explicit verbalization of evaluation awareness. Our results demonstrate that evaluation meta-knowledge may inflate safety benchmark performance, introducing a novel confounder that is independent of explicit memorization or verbalized evaluation awareness, thus, challenging to detect. These findings have important implications for the design and interpretation of AI safety evaluations. Our code and models are available at https://github.com/compass-group-tue/arxiv2026 evaluation meta knowledge.
Sources
- Taken out of context: On measuring situational awareness in LLMs
- Evaluation Faking: Unveiling Observer Effects in Safety Evaluation of Frontier AI Systems
- GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models
- Alignment faking in large language models
- Agentic Misalignment: How LLMs Could Be Insider Threats
- Natural Emergent Misalignment from Reward Hacking in Production RL
- Auditing language models for hidden objectives
- Qwen3 Technical Report
- Stress Testing Deliberative Alignment for Anti-Scheming Training
- AuditBench: Evaluating Alignment Auditing Techniques on Models with Hidden Behaviors
- Believe It or Not: How Deeply do LLMs Believe Implanted Facts?
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering