The Optimizer Is the Agent: Reasoning-Driven Search across Prompts, Programs, and ML Workflows
cs.AI
Submitted: 2026-08-07
Updated: 2026-08-31
Journal ref: COLM 2026
Code: https://github.com/deepseek-ai/eplb
Project page: https://erich-friedman.github.io/packing/cirRsqu
License: http://creativecommons.org/licenses/by/4.0/
The gist: Recent systems for optimizing prompts, programs, and ML workflows typically rely on explicit outer-loop controllers such as evolutionary search, bandits, or textual-gradient methods.
Terminology
Abstract
Recent systems for optimizing prompts, programs, and ML workflows typically rely on explicit outer-loop controllers such as evolutionary search, bandits, or textual-gradient methods. We ask a fundamentally different question: how much of this search policy can be internalized by a single tool-using agent? We present ReASearch, a unified framework for reasoning-driven optimization in which the agent autonomously decides what to evaluate, how to diagnose failures, which edits to make, and when to verify or restart. Rather than serving only as a proposal generator guided by hand-designed heuristics, the agent actively analyzes outcomes, allocates budget, and refines its strategy over long horizons through persistent memory. With a shared agent loop and domain-specific tools, ReASearch instantiates the exact same scaffold to optimize prompts, programs, and ML workflows. Across 14 diverse tasks, it is competitive with and mostly better than specialized optimization systems, achieving gains of 2% to 40% over strong domain-specific baselines, and in some cases discovering solutions that improve on prior human best-known results. Crucially, we observe that complex search behaviors, which are typically implemented by explicit controllers, emerge naturally from the agent's reasoning process.
Sources
- GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning
- CodeEvolve: an open source evolutionary coding agent for algorithmic discovery and optimization
- AdaEvolve: Adaptive LLM Driven Zeroth-Order Optimization
- ARC-AGI-2: A New Challenge for Frontier AI Reasoning Systems
- Training Verifiers to Solve Math Word Problems
- Scaling Textual Gradients via Sampling-Based Momentum
- Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution
- DeltaEvolve: Accelerating Scientific Discovery through Momentum-Driven Evolution
- DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines
- ShinkaEvolve: Towards Open-Ended And Sample-Efficient Program Evolution
- Efficient Multi-objective Prompt Optimization via Pure-exploration Bandits
- Agentic Harness Engineering: Observability-Driven Automatic Evolution of Coding-Agent Harnesses
- EvoX: Meta-Evolution for Automated Discovery
- Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
- AlphaEvolve: A coding agent for scientific and algorithmic discovery
- Hyperband-based Bayesian Optimization for Black-box Prompt Selection
- From Computational Certification to Exact Coordinates: Heilbronn's Triangle Problem on the Unit Square Using Mixed-Integer Optimization
- PromptAgent: Strategic Planning with Language Models Enables Expert-level Prompt Optimization
- LLM Prompt Duel Optimizer: Efficient Label-Free Prompt Optimization
- PACEvolve: Enabling Progress-Aware Consistent Evolution
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection