Evo-Harness: Context-to-Harness Skill Compilation for Self-Evolving Agents
cs.AI, cs.CL
Submitted: 2026-08-15
Updated: 2026-08-30
Comments: EMNLP 2026 Main
Code: https://github.com/A-EVO-Lab/a-evolve
License: http://creativecommons.org/licenses/by/4.0/
The gist: Learning from experience is critical for developing capable, self-improving large language model (LLM) agents.
Terminology
Abstract
Learning from experience is critical for developing capable, self-improving large language model (LLM) agents. Existing methods typically extract knowledge from accumulated trajectories via reflection, memory, rules, or skills. However, agents in realistic environments continuously encounter novel tasks, often offering only a one-shot opportunity to improve. These executions yield rich but highly noisy contexts, entangling broadly useful lessons with task-specific artifacts. Critically, prior works rarely validate their effectiveness on complex real-world tasks or isolate the underlying drivers of improvement. To address these gaps, we formulate online harness learning, where a frozen agent improves by continually updating a structured harness across sequential tasks. This formulation enables a systematic study of key self-improvement factors through our proposed Evo-Harness. At its core, context-to-harness skill compilation distills noisy, single-shot executions into reusable skill harnesses for cross-domain and topic-level adaptation. To demonstrate the efficacy of one-shot skill compilation, we evaluate across five realistic benchmarks (TerminalBench2, SWE-bench, CL-Bench, -bench, WebArena-Infinity). Our extensive analysis demonstrates the effectiveness of Evo-Harness and provides a principled understanding of how LLM agents can effectively learn on the fly. Our code is available at https://github.com/A-EVO-Lab/a-evolve/tree/release/evo-harness.
Sources
- Plan-and-Act: Improving Planning of Agents for Long-Horizon Tasks
- A Survey of Agent Memory in the Second Half: Towards Self-Evolving and Long-Horizon Agents
- XSkill: Continual Learning from Experience and Skills in Multimodal Agents
- GPT-4 Technical Report
- gpt-oss-120b & gpt-oss-20b Model Card
- Meta-Harness: End-to-End Optimization of Model Harnesses
- GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning
- Mem-Gallery: Benchmarking Multimodal Long-Term Conversational Memory for MLLM Agents
- Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory
- CL-bench: A Benchmark for Context Learning
- Position: Agentic Evolution is the Path to Evolving LLMs
- Harness Updating Is Not Harness Benefit: Disentangling Evolution Capabilities in Self-Evolving LLM Agents
- DeepSeek-V3 Technical Report
- Omni-SimpleMem: Autoresearch-Guided Discovery of Lifelong Multimodal Agent Memory
- Adaptive Auto-Harness: Sustained Self-Improvement for Agentic System Deployment on Open-Ended Task Streams
- AutoHarness: improving LLM agents by automatically synthesizing a code harness
- Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
- Trace2Skill: Distill Trajectory-Local Lessons into Transferable Agent Skills
- Code as Agent Harness
- SkillOS: Learning Skill Curation for Self-Evolving Agents
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection