A Theory of Reliable Self-Evolution for Agent Harnesses
cs.AI
Submitted: 2026-09-08
Updated: 2026-09-28
Code: https://github.com/deepseek-ai/deepseek-harness
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- MOSS: Self-Evolution through Source-Level Rewriting in Autonomous Agent Systems
- Evaluating Large Language Models Trained on Code
- HarnessX: A Composable, Adaptive, and Evolvable Agent Harness Foundry
- Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution
- Oracle Gap and Signal Fidelity: A Fixed-Pool Diagnostic for Test-Time Collaboration
- Meta-Harness: End-to-End Optimization of Model Harnesses
- Reinforcement Learning and Control as Probabilistic Inference: Tutorial and Review
- Agentic Harness Engineering: Observability-Driven Automatic Evolution of Coding-Agent Harnesses
- Information Theoretic Guarantees For Policy Alignment In Large Language Models
- TTHE: Test-Time Harness Evolution
- AlphaEvolve: A coding agent for scientific and algorithmic discovery
- A Self-Improving Coding Agent
- Self-Evolving Software Agents
- Information-Theoretic Limits of Safety Verification for Self-Improving Systems
- AgentSquare: Automatic LLM Agent Search in Modular Design Space
- Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
- Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models
- Learning from Failure: Inference-Time Self-Improvement for Computer-Use Agents
- On The Statistical Limits of Self-Improving Agents
- Huxley-G\"odel Machine: Human-Level Coding Agent Development by an Approximation of the Optimal Self-Improving Machine
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection