EvoHarness-RL: Learning Runtime Harness Coordination for Self-Evolving Agents
cs.LG, cs.CL
Submitted: 2026-08-05
Updated: 2026-10-05
Comments: Accepted to LLA@COLM 2026
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: Long-horizon LLM agents increasingly rely on external execution support to maintain state, track progress, invoke tools, verify outcomes, and reuse experience across interactions.
Terminology
Abstract
Long-horizon LLM agents increasingly rely on external execution support to maintain state, track progress, invoke tools, verify outcomes, and reuse experience across interactions. However, effective harness use raises two coupled challenges: state formation from noisy interaction traces and runtime control over external-state access. Existing agents usually handle both through prompts, heuristics, or domain-specific conventions, leaving the external workspace and its usage policy manually engineered. To address this, we study the problem of harness policy learning, where agents learn harness policies offline and deploy them to construct and update external harness state online during runtime task execution. We introduce EvoHarness-RL, which exposes Belief, Progress, and Experience (BPE) as policy-facing harness state. Supervised harness fine-tuning teaches the base agent the harness action space and how to construct useful external state, while cost-aware GRPO explores coordination policies to selectively read, update, and consolidate that state during long-horizon interaction. Instantiated on ALFWorld with a Qwen3-8B LLM, EvoHarness-RL reaches 96.9% success and reveals two key dynamics: harness annealing, where training internalizes recurring harness-use patterns into the model policy and shifts the agent from frequent harness calls toward selective external-state access, and harness evolution, where progress updates and experience consolidation refine the harness into a compact, task-adaptive state substrate. These results suggest that long-horizon agents benefit from trainable policies for constructing and coordinating with external harness workspaces, beyond simply adding stronger tools or larger memories.
Sources
- HarnessX: A Composable, Adaptive, and Evolvable Agent Harness Foundry
- Harness-1: Reinforcement Learning for Search Agents with State-Externalizing Harnesses
- Meta-Harness: End-to-End Optimization of Model Harnesses
- AutoHarness: improving LLM agents by automatically synthesizing a code harness
- Code as Agent Harness
- ReasoningBank: Scaling Agent Self-Evolving with Reasoning Memory
- SkillOS: Learning Skill Curation for Self-Evolving Agents
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Agent-BRACE: Decoupling Beliefs from Actions in Long-Horizon Tasks via Verbalized State Uncertainty
- LiveResearchBench: A Live Benchmark for User-Centric Deep Research in the Wild
- A Subgoal-driven Framework for Improving Long-Horizon LLM Agents
- Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory
- SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning
- ReAct: Synergizing Reasoning and Acting in Language Models
- Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks