Self-Healing Harness for Runtime Oversight of Agent Self-Modification
cs.AI
Submitted: 2026-09-21
Updated: 2026-09-21
License: http://creativecommons.org/licenses/by/4.0/
The gist: LLM agents can change their own future behavior, raising a basic control question of which self-generated changes should be allowed to persist.
Terminology
Abstract
LLM agents can change their own future behavior, raising a basic control question of which self-generated changes should be allowed to persist. We formulate this as admission control for self-modification. The agent may propose changes to its operating instructions, while an external runtime gate controls persistence. We implement this principle as a model-agnostic self-healing harness that runs a Detect, Notice, Heal, Validate loop around an otherwise unmodified agent. The agent authors candidate behavioral rules in an external workspace, where they receive provisional execution authority during evaluation and acquire persistent cross-episode authority only after measured improvement on the triggering failure without regression beyond a fixed margin on protected cases. Replay provides matched evidence when available, forward trials provide a weaker fallback, and a corpus-level guard re-tests the accumulated active rule set. Across 16 matched Baseline and Harness runs spanning AppWorld, Terminal-Bench, and τ squared-Bench, the gate rejected 383 replay-decided proposals. Of these, 211 (55%) improved their triggering failure while degrading a case that previously worked. This shows that locally beneficial self-modifications can introduce collateral regressions often enough to materially affect gate decisions, providing direct empirical motivation for external admission control. Task-completion score is higher under the Harness in all 16 pairs, with two paired bootstrap intervals excluding zero, while repeated-trial reliability is higher in 12 pairs, tied in 4, and lower in none. Because adaptation modifies the policy-inducing context while leaving model weights fixed, admitted changes remain inspectable, reversible, and compatible with closed-weight models.
Sources
- Concrete Problems in AI Safety
- Constitutional AI: Harmlessness from AI Feedback
- $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment
- Measuring Progress on Scalable Oversight for Large Language Models
- Learnable Conformal Prediction with Context-Aware Nonconformity Functions for Robotic Planning and Perception
- MemGPT: Towards LLMs as Operating Systems
- TRACER: Trajectory Risk Aggregation for Critical Episodes in Agentic Reasoning
- Learning Conformal Abstention Policies for Adaptive Risk Management in Large Language and Vision-Language Models
- Voyager: An Open-Ended Embodied Agent with Large Language Models
- Agent Workflow Memory
- A-MEM: Agentic Memory for LLM Agents
- $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
- Self-Taught Optimizer (STOP): Recursively Self-Improving Code Generation
- Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection