StarHarness: Evolving Harnesses with Stratified Search for Enterprise Environments

arXiv:2608.24804 · cs.AI, cs.SE · Submitted 2026-08-25 · Read on arXiv

cs.AI, cs.SE

Submitted: 2026-08-25

Updated: 2026-08-25

Code: https://github.com/ServiceNow/StarHarness

License: http://creativecommons.org/licenses/by/4.0/

The gist: We present StarHarness, a framework for evolving environment-specific agent harnesses while keeping model weights fixed.

Terminology

Abstract

We present StarHarness, a framework for evolving environment-specific agent harnesses while keeping model weights fixed. The evolved harness can include prompt and task framing, tool interfaces, skills, MCP-backed providers, subagent structure, and agent-loop configuration. StarHarness constructs a compact evolution pool by stratifying tasks according to baseline failure behavior, separates proposer-visible search tasks from proposer-hidden selection tasks, and reserves held-out tasks for evaluating generalization. Across ITBench SRE, EnterpriseOps-Gym ITSM, and AutomationBench Finance, harness evolution improves full-benchmark performance by 20-35 percentage points over the default harness after 4-12 accepted changes per environment. These gains persist on tasks excluded from evolution and transfer without re-evolution across GPT and Qwen model families. Trace analysis links the improvements to interface repairs, environment conventions, and operational knowledge that compresses search, with fewer false-positive diagnoses and shorter trajectories in several settings. StarHarness therefore offers a practical way to reduce persistent model-environment mismatch in tool-rich enterprise tasks.

Related papers