Competing at Every Price Point with Agentic Evolution over a Menu of LLMs
cs.AI
Submitted: 2026-08-17
Updated: 2026-09-08
Comments: Code at https://github.com/andborth/RoboPhD
Code: https://github.com/andborth/RoboPhD
License: http://creativecommons.org/licenses/by/4.0/
The gist: Consider a firm that surveys its competition for a particular agentic task and seeks to offer superior accuracy at every price point.
Terminology
Abstract
Consider a firm that surveys its competition for a particular agentic task and seeks to offer superior accuracy at every price point. A firm that Pareto-dominated its competitors would leave no rational customer a reason to buy elsewhere. This paper shows a path to this kind of capability by evolving multi-LLM Python agents from training pools of at most 100 examples. Given a priced menu of nine LLM endpoints; brief documentation of the task, objective, and API; a simple seed agent; and an operator-chosen per-problem cost target--usually set at an incumbent's own price--RoboPhD, an evolutionary meta-agent, evolves complete agent programs that attack the public frontiers of two semantically dissimilar tasks point by point: DS-1000 (execution-checked code generation) and PaperFindingBench (LLM-judged scientific document retrieval). On public leaderboards for each task, the evolved agents hold every Pareto-frontier slot but one, including Pareto domination of both the top-scoring and the lowest-cost competing points.
Sources
- GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning
- Concrete Problems in AI Safety
- RoboPhD: Evolving Diverse Complex Agents Under Tight Evaluation Budgets
- FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance
- Optimizing Model Selection for Compound AI Systems
- HarnessForge: Joint Harness and Policy Evolution for Adaptive Agent Systems
- The Semantic Scholar Open Data Platform
- Meta-Harness: End-to-End Optimization of Model Harnesses
- The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery
- EvoScientist: Towards Multi-Agent Evolving AI Scientists for End-to-End Scientific Discovery
- AlphaEvolve: A coding agent for scientific and algorithmic discovery
- RouteLLM: Learning to Route LLMs with Preference Data
- HARBOR: Automated Harness Optimization
- Conformal Constrained Policy Optimization for Cost-Effective LLM Agents
- Rethinking the Evaluation of Harness Evolution for Agents
- Self-Harness: Harnesses That Improve Themselves
- Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection