Auditing Harness Tampering in Self-Improving Agents
cs.CL, cs.AI
Submitted: 2026-08-30
Updated: 2026-08-30
License: http://creativecommons.org/licenses/by/4.0/
The gist: Self-improving agents iteratively modify their own harness to push the frontier of their performance.
Terminology
Abstract
Self-improving agents iteratively modify their own harness to push the frontier of their performance. However, such modifications can produce illusory performance gains or compromise integrity constraints such as authorization, provenance, and completeness without genuinely improving capability. We term this phenomenon as harness tampering, which extends the concept from reward and measurement tampering to the full self-improvement lifecycle. To systematically study this problem, we propose a two-axis taxonomy that categorizes each misaligned edit by the harness functional role in which it occurs and the obligation it violates. Then we build an annotated corpus by seeding tampered-benign edit pairs into the real trajectories of self-improving agents. We adapt and benchmark diverse audit methods on tampering classification and localization tasks. Finally we systematically audit real trajectories of self-improving agents. The results demonstrate that harness tampering consistently occurs in real runs from different agents, often persists in the lineage of the best agent, and forms distinct system-specific profiles across the taxonomy.
Sources
- Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
- Benchmarking Reward Hack Detection in Code Environments via Contrastive Analysis
- A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems
- Distributed Attacks in Persistent-State AI Control
- SHADE-Arena: Evaluating Sabotage and Monitoring in LLM Agents
- Meta-Harness: End-to-End Optimization of Model Harnesses
- ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D
- Harness Updating Is Not Harness Benefit: Disentangling Evolution Capabilities in Self-Evolving LLM Agents
- Safety in Self-Evolving LLM Agent Systems: Threats, Amplification, and Case Studies
- Harness-Aware Self-Evolving: Co-Evolving Model Weights, Harness, and Task Solutions
- HarnessBank: Semantic Gene-Bank Search with Gated Verification for Agent-Harness Self-Evolution
- ScientistOne: Towards Human-Level Autonomous Research via Chain-of-Evidence
- Code as Agent Harness
- Self-Improvements in Modern Agentic Systems: A Survey
- Benchmarks for Detecting Measurement Tampering
- Hack-Verifiable Environments: Towards Evaluating Reward Hacking at Scale
- Reward Hacking Benchmark: Measuring Exploits in LLM Agents with Tool Use
- Who Guards the Benchmarks? Automated Auditing of LLM Agent Benchmarks
- Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack
- Rethinking the Evaluation of Harness Evolution for Agents
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering