SpecBench: Measuring Reward Hacking in Long-Horizon Coding Agents
cs.SE, cs.AI, cs.CL
Submitted: 2026-05-20
Updated: 2026-09-09
Code: https://github.com/anomalyco/opencode
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Program Synthesis with Large Language Models
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation
- Terminal Wrench: A Dataset of 331 Reward-Hackable Environments and 3,632 Exploit Trajectories
- Qwen3-Coder-Next Technical Report
- DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
- Evaluating Large Language Models Trained on Code
- SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
- Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
- Benchmarking Reward Hack Detection in Code Environments via Contrastive Analysis
- NL2Repo-Bench: Towards Long-Horizon Repository Generation Evaluation of Coding Agents
- EvilGenie: A Reward Hacking Benchmark
- Scaling Laws for Reward Model Overoptimization
- DevBench: A Realistic, Developer-Informed Benchmark for Code Generation Models
- Alignment faking in large language models
- LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
- Countdown-Code: A Testbed for Studying The Emergence and Generalization of Reward Hacking in RLVR
- Kimi K2.5: Visual Agentic Intelligence
- AIDE: AI-Driven Exploration in the Space of Code
- Categorizing Variants of Goodhart's Law
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties