Monitoring and Discovering Reward Hacking with Internal Representations during LLM Evaluations
cs.CL, cs.LG
Submitted: 2026-09-16
Updated: 2026-09-16
License: http://creativecommons.org/licenses/by/4.0/
The gist: As models scale, reward hacking becomes more frequent, more sophisticated, and more consequential.
Terminology
Abstract
As models scale, reward hacking becomes more frequent, more sophisticated, and more consequential. Does it leave a telltale signature in model representations? This work analyzes how reward hacking is represented internally in frontier open source LLMs, and how those representations can be used to understand and discover the range of hacking behaviors a model displays. In particular, we find that simple difference of means vectors coherently represent reward hacking in Kimi K3, GLM 5.2, and Qwen 3.8 Max across a variety of behaviors in common evaluations. Despite their simplicity, these vectors are both generalizable and interpretable, and we can use them to reliably detect reward hacking. We first evaluate reward hacking in commonly reported benchmarks like DeepSWE and SWE-bench, finding that models reward hack excessively in these environments; GLM 5.2 hacks in 57.2% of rollouts on DeepSWE and in 73% of rollouts on SWE-bench. Catching these requires monitors; LLM monitors are effective, but expensive detectors. We show that DoM vectors are similarly effective but virtually free, catching 3.1% more hacks in Kimi K3 and 7.9% fewer hacks in GLM 5.2 on DeepSWE at a monitor matched false positive rate. DoM vectors run on the chain-of-thought also predict reward hacks in the model's subsequent actions, meaning we can run them online and catch potential hacks before they occur. Finally, we analyze probe-hits that LLM monitors do not catch and discover other undesirable behaviors, as well as show transfer to finding hacks in non-SWE evaluations. Together, these results provide evidence that simple, white-box methods can be used to scalably study and monitor reward hacking behaviors in frontier open source models
Sources
- Concrete Problems in AI Safety
- Constitutional AI: Harmlessness from AI Feedback
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation
- $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment
- Terminal Wrench: A Dataset of 331 Reward-Hackable Environments and 3,632 Exploit Trajectories
- Anatomy of Post-Training: Using Interpretability to Characterize Data and Shape the Learning Signal
- Demonstrating specification gaming in reasoning models
- MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
- Persona Vectors: Monitoring and Controlling Character Traits in Language Models
- Models That Know How Evaluations Are Designed Score Safer
- Implicit Chain of Thought Reasoning via Knowledge Distillation
- Output Supervision Can Obfuscate the Chain of Thought
- Concise Reasoning via Reinforcement Learning
- EvilGenie: A Reward Hacking Benchmark
- DeepSWE: Measuring Frontier Coding Agents on Original, Long-Horizon Engineering Tasks
- The Art of Scaling Reinforcement Learning Compute for LLMs
- Reinforcement Learning from Human Feedback
- Decomposing and Measuring Evaluation Awareness
- TACO: Topics in Algorithmic COde generation dataset
- ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering