Rethinking Self-Evolving Agents: Do We Still Need Prescribed Optimization Pipelines?
Hui Xue, Fan Yang
Microsoft Research
cs.AI
Submitted: 2026-08-10
Updated: 2026-08-11
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 75/100
The gist: This paper investigates whether self-evolving agents require prescribed optimization pipelines—frameworks that determine how to gather evidence, revise persistent artifacts, select candidates, and
Terminology
Summary
This paper investigates whether self-evolving agents require prescribed optimization pipelines—frameworks that determine how to gather evidence, revise persistent artifacts, select candidates, and stop—or whether a sufficiently capable frontier model can compose the improvement process itself online. The authors ask: Under the same external contract, must the framework prescribe the task-specific optimization meta-policy, or can a capable optimizer compose it online?
The paper introduces a distinction between two components of self-evolving systems:
-
The optimization contract: Fixed elements including
the objective, permitted interactions, resource budget, data boundary, and evaluation
that the framework must always govern. -
The optimization meta-policy:
The task-specific logic that organizes evidence, revision, selection, and stopping
—the route from evidence to persistent improvement.
The authors introduce Open-Ended Optimization (OEO) as a protocol that "preserves the external contract and permitted task-facing operations, but does not impose a reflection template, edit language, candidate schedule, or task-specific stopping rule. Instead, the optimizer composes the improvement process online."
The paper compares OEO against two prescribed approaches with deliberately different search biases:
-
SKILLOPT:
a staged, bounded skill-training pipeline
with fixed rollout/reflection batches, success/failure diagnosis separation, bounded patch operators, and validation-gated updates -
GEPA:
a reflective evolutionary search
with free-form textual mutation inside a Pareto-aware evolutionary loop
All methods start from the same initial skill and operate under the same optimization contract across 8 benchmark–target-model settings (SearchQA, SpreadsheetBench, OfficeQA, and LiveMath, with Qwen3.5-4B and GPT-5.5 as target models).
With GPT-5.5 as the optimizer, OEO records 7 wins and 1 tie against SKILLOPT across 8 benchmark–target-model settings, and 5 wins against GEPA in 6 confirmatory settings; GEPA's only lead is 0.21 percentage points.
OEO improves every initial skill across all 8 settings. Notably, OEO uses a median 34.3% of the configured SKILLOPT target-interaction token budget
and stays below the reference budget in every setting, demonstrating the result is not explained by spending more target-interaction tokens.
A static-input-matched, zero-interaction rewrite control shows that a single static frontier-model rewrite does not reproduce OEO in any tested cell.
While the one-shot rewrite helps on SearchQA (improving by 2.07 and 5.00 percentage points), it still remains 6.14 and 3.57 percentage points below OEO.
On LiveMath, the control actually lowers both initial scores, while OEO improves them by 34.68 and 29.84 percentage points.
The paper identifies a clear capability boundary for delegation. With protocols frozen, SKILLOPT outperforms OEO at medium optimizer capability on both ladder tasks
(leading by 9.68 percentage points on LiveMath and 3.50 percentage points on SearchQA). A weak optimizer cannot produce a valid action through the unchanged OEO interface
(blocked), while the SKILLOPT runner completes both runs. Moving from medium to frontier optimizer raises OEO by 16.13 percentage points on LiveMath and 7.07 percentage points on SearchQA, compared with only 2.42 and 0.64 percentage points for SKILLOPT, showing online meta-policy composition depends directly on optimizer capability.
In the fully instrumented OEO–SKILLOPT pair, trajectory analysis reveals that prescription changes the optimization path more consistently than it changes the selected skill's evaluated behavior.
Across all 8 settings, OEO makes a larger maximum edit, touches a broader set of sections, and exhibits greater revision churn than SKILLOPT.
Yet pass/fail agreement is at least 0.78 in 7 of 8 settings, and correct-set Jaccard exceeds 0.70 in 5.
The paper notes that different procedures can traverse substantially different paths yet reach overlapping item-level behavior, while similar aggregate scores can conceal complementary correct-item sets.
The authors conclude that prescribed pipelines are capability-dependent scaffolding: essential constraints remain external, but a sufficiently capable optimizer can compose the route from measurable feedback to persistent improvement.
They propose a capability-adaptive division of labor
where the framework should continue to own objectives, permissions, budgets, evaluation, and governance, while delegating the task-specific route to improvement when the model can carry it.
Prescription is neither universally necessary nor uniformly redundant: it can scaffold optimization when the model cannot reliably compose the improvement process for itself.
Improvements for AI systems
Improvements to AI Systems:
-
Add capability-adaptive optimization delegation. AI systems can dynamically assess their own optimizer capability (e.g., via self-evaluation on small validation tasks) and choose between a prescribed pipeline (like SKILLOPT) and open-ended self-improvement (OEO). If capability is low, the system falls back to structured scaffolding; if high, it composes its own improvement strategy online. This yields more robust performance across heterogeneous tasks and model strengths.
-
Implement budget-aware self-improvement. Systems can monitor their own target-interaction token usage in real time and terminate or re-plan improvement steps when they exceed a fraction (e.g., 34.3%) of the configured budget, as OEO does. This reduces compute cost and prevents over-optimization, while still achieving superior or equal results to fixed-budget pipelines.
-
Enable online meta-policy composition without fixed templates. Replace rigid reflection templates, edit languages, and candidate schedules with a single, flexible interface that lets the optimizer decide how to gather evidence, revise artifacts, select candidates, and stop. This allows the system to adapt its improvement process to the specific task, leading to higher gains (e.g., +34.68 percentage points on LiveMath) than any static pipeline.
-
Add one-shot rewrite rejection for safety-critical tasks. Before applying a static frontier-model rewrite, the system can run a small pilot evaluation on a held-out subset. If the rewrite lowers performance (as seen on LiveMath), the system discards it and switches to iterative, feedback-driven improvement. This prevents regressions from single-shot interventions.
-
Use process-outcome divergence detection for better ensemble selection. When multiple optimization paths (e.g., OEO vs. SKILLOPT) produce similar aggregate scores but different correct-item sets (Jaccard > 0.70 but not identical), the system can merge or ensemble the outputs from both paths to maximize coverage. This improves recall on diverse question types without additional training.
-
Deploy capability-gated delegation in multi-agent systems. In a team of agents with different underlying models, the system can route optimization tasks to the most capable agent for open-ended improvement, while less capable agents use prescribed pipelines. This yields higher overall performance and avoids invalid actions from weaker optimizers, as shown by the capability boundary results.
-
Build self-monitoring for revision churn. The system can track edit size, section breadth, and revision frequency during self-improvement. If churn exceeds a threshold without pass/fail improvement, it can switch to a more conservative, validation-gated update rule (like SKILLOPT) to avoid destabilizing the artifact. This balances exploration and stability.
-
Enable cross-task transfer of composed meta-policies. Since OEO composes improvement strategies online, the system can store successful meta-policies (e.g., “first diagnose failures, then mutate only the relevant section, then validate”) as reusable templates for similar future tasks. This reduces the need for re-composition and speeds up adaptation in new domains.
What the improved AI system can do:
-
Self-improve on new benchmarks (SearchQA, SpreadsheetBench, OfficeQA, LiveMath) with higher accuracy than any fixed pipeline, using less compute.
-
Automatically switch between structured and open-ended optimization based on its own capability, avoiding failures from weak optimizers.
-
Reject harmful one-shot rewrites and recover from regressions.
-
Combine outputs from different improvement paths to answer a broader set of correct items.
-
Operate efficiently under token budgets, making it suitable for cost-sensitive deployment.
-
Transfer learned improvement strategies across tasks, reducing time-to-solution for novel problems.
Abstract
Self-evolving agents are usually built around prescribed optimization pipelines: the framework decides how to gather evidence, revise a persistent artifact, select candidates, and stop. We ask whether this task-specific procedure remains necessary when a frontier model acts as the optimizer. We introduce Open-Ended Optimization (OEO), which keeps the objective, permitted interactions, resource budget, data boundary, and evaluation fixed while allowing the optimizer to compose the improvement process online. We compare OEO with two complementary prescribed approaches: SkillOpt, a staged pipeline with bounded edits, and GEPA, a reflective evolutionary search. Across 14 head-to-head comparisons over 8 benchmark-target-model settings, GPT-5.5-driven OEO records 12 wins, 1 tie, and 1 narrow loss of 0.21 percentage points. It uses a median 34.3 percent of SkillOpt's configured target-interaction token budget. A one-shot, zero-interaction control shows that the gains are not explained by a single prior-driven rewrite. However, delegation has a capability boundary: SkillOpt outperforms OEO with a medium optimizer, and a weak optimizer cannot operate through the unchanged OEO interface. In the fully instrumented OEO-SkillOpt pair, trajectory analysis further shows that prescription changes how optimization proceeds more consistently than it changes final behavior. Together, these findings recast prescribed pipelines as capability-dependent scaffolding: essential constraints remain external, but a sufficiently capable optimizer can compose the route from measurable feedback to persistent improvement.
Sources
- Voyager: An Open-Ended Embodied Agent with Large Language Models
- Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution
- TextGrad: Automatic "Differentiation" via Text
- Automated Design of Agentic Systems
- AFlow: Automating Agentic Workflow Generation
- Harnessing Agentic Evolution
- Bayesian-Agent: Posterior-Guided Skill Evolution Across LLM Agent Harnesses
- SkillOS: Learning Skill Curation for Self-Evolving Agents
- Hypothesis-Driven Skill Optimization for LLM Agents
- SkillAudit: Ground-Truth-Free Skill Evolution via Paired Trajectory Auditing
- SkillOpt: Executive Strategy for Self-Evolving Agent Skills
- SkillOpt-Lite: Better and Faster Agent Self-evolution via One Line of Vibe
- SearchQA: A New Q&A Dataset Augmented with Context from a Search Engine
- SpreadsheetBench: Towards Challenging Real World Spreadsheet Manipulation
- OfficeQA Pro: An Enterprise Benchmark for End-to-End Grounded Reasoning
- LiveMathematicianBench: A Live Benchmark for Research-Level Mathematical Reasoning with Proof Sketches
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection