Countdown-Code: A Testbed for Studying The Emergence and Generalization of Reward Hacking in RLVR
cs.LG, cs.AI, cs.CL
Submitted: 2026-03-07
Updated: 2026-09-11
Comments: EMNLP 2026 camera ready
Code: https://github.com/zohaib-khan5040/Countdown-Code
License: http://creativecommons.org/licenses/by/4.0/
The gist: Reward hacking is a form of misalignment in which models overoptimize proxy rewards without genuinely solving the underlying task.
Terminology
Abstract
Reward hacking is a form of misalignment in which models overoptimize proxy rewards without genuinely solving the underlying task. Precisely measuring reward hacking occurrence remains challenging because true task rewards are often expensive or impossible to compute. We introduce Countdown-Code, a minimal environment where models can both solve a mathematical reasoning task and manipulate the test harness. This dual-access design creates a clean separation between proxy rewards (test pass/fail) and true rewards (mathematical correctness), enabling accurate measurement of reward-hacking rates. Using this environment, we study reward hacking in open-weight LLMs and find that such behaviors can be unintentionally learned during supervised fine-tuning (SFT) when even a small fraction of reward-hacking trajectories leak into training data. As little as 1% contamination in distillation SFT data is sufficient for models to internalize reward hacking which resurfaces during subsequent reinforcement learning (RL). We further show that RL amplifies misalignment and drives its generalization beyond the original domain. We open-source our environment and code to facilitate future research on reward hacking in LLMs. Our results reveal a previously underexplored pathway through which reward hacking can emerge and persist in LLMs, underscoring the need for more rigorous validation of synthetic SFT data. Code is available at https://github.com/zohaib-khan5040/Countdown-Code.
Sources
- Concrete Problems in AI Safety
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation
- Evaluating Large Language Models Trained on Code
- Reasoning Models Don't Always Say What They Think
- SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training
- Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs
- Monitoring Monitorability
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- OpenAI o1 System Card
- Goodhart's Law in Reinforcement Learning
- Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
- LLMs Can Easily Learn to Reason from Demonstrations Structure, not content, is what matters!
- The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Defining and Characterizing Reward Hacking
- Poisoning Attacks on LLMs Require a Near-constant Number of Poison Samples
- Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting
- Is It Thinking or Cheating? Detecting Implicit Reward Hacking by Measuring Reasoning Effort
- Inoculation Prompting: Instructing LLMs to misbehave at train-time improves test-time alignment
- LeetCodeDataset: A Temporal Dataset for Robust Evaluation and Efficient Training of Code LLMs
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks