Training LLMs to Verbalize Evaluation Awareness
cs.CL, cs.AI, cs.LG
Submitted: 2026-09-28
Updated: 2026-09-28
Code: https://github.com/tim-hua-01/opus-4-6-openrouter-weirdness
Project page: https://elena-baixy.github.io/verbalization.html
Terminology
Sources
- AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
- Foundational Challenges in Assuring Alignment and Safety of Large Language Models
- Analyzing and Improving Chain-of-Thought Monitorability Through Information Theory
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation
- Taken out of context: On measuring situational awareness in LLMs
- Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs
- Discovering Latent Knowledge in Language Models Without Supervision
- Models That Know How Evaluations Are Designed Score Safer
- Steering Evaluation-Aware Language Models to Act Like They Are Deployed
- Evaluation Awareness in Language Models Has Limited Effect on Behaviour
- Training Agents to Self-Report Misbehavior
- Decomposing and Measuring Evaluation Awareness
- Agentic Misalignment: How LLMs Could Be Insider Threats
- Eliciting Latent Knowledge from Quirky Language Models
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
- Large Language Models Often Know When They Are Being Evaluated
- Probing and Steering Evaluation Awareness of Language Models
- Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
- CIAware-Bench: Benchmarking Control Intervention Awareness Across Frontier LLMs
- Stress Testing Deliberative Alignment for Anti-Scheming Training
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering