Compute Allocation for Self-Evolving LLMs: From Depth-Breadth to Multi-Armed Bandits
cs.CL, cs.AI, cs.LG, cs.NE
Submitted: 2026-05-28
Updated: 2026-09-12
Code: https://github.com/keruiwu/self-evolving-allocation
License: http://creativecommons.org/licenses/by/4.0/
The gist: LLM-guided evolutionary search (Evolve systems) has reached state-of-the-art results on mathematical and combinatorial tasks, yet most existing systems report only the best of many runs and leave the
Terminology
Abstract
LLM-guided evolutionary search (Evolve systems) has reached state-of-the-art results on mathematical and combinatorial tasks, yet most existing systems report only the best of many runs and leave the run-to-run distribution undocumented. We ask how a fixed budget of LLM calls should be allocated, and how reliably a single run reaches the reported numbers. Sweeping the depth-breadth grid over five models and three tasks, we identify two empirical regularities: a fitness-compute envelope along which capability ordering largely collapses when measured in effective FLOPs, and a bilinear depth-breadth fit with task-specific interaction; both are gated by model-task capability. Motivated by these regularities, we propose BaSE (Bandit-based Self-Evolving), a multi-armed bandit that allocates LLM calls across parallel trajectories. Without changing the model, prompt, or evaluator, BaSE improves mean fitness by 12.3% over the strongest island-protocol baseline across 8 (model, task) cells, with the largest gains on high-variance settings: a reliability gain from allocation alone.
Sources
- Deep Reinforcement Learning at the Edge of the Statistical Precipice
- Training Compute-Optimal Large Language Models
- CodeEvolve: an open source evolutionary coding agent for algorithmic discovery and optimization
- Wider or Deeper? Scaling LLM Inference-Time Compute with Adaptive Branching Tree Search
- Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
- AdaEvolve: Adaptive LLM Driven Zeroth-Order Optimization
- EvoPrompting: Language Models for Code-Level Neural Architecture Search
- ShinkaEvolve: Towards Open-Ended And Sample-Efficient Program Evolution
- Provable Scaling Laws for the Test-Time Compute of Large Language Models
- Evolving Deeper LLM Thinking
- Evolution through Large Models
- ThetaEvolve: Test-time Learning on Open Problems
- Aligning Tree-Search Policies with Fixed Token Budgets in Test-Time Scaling of LLMs
- ParaThinker: Native Parallel Thinking as a New Paradigm to Scale LLM Test-time Compute
- Illuminating search spaces by mapping elites
- Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models
- AlphaEvolve: A coding agent for scientific and algorithmic discovery
- Large Language Models as Optimizers
- Population-Evolve: a Parallel Sampling and Evolutionary Method for LLM Math Reasoning
- The Sequential Edge: Inverse-Entropy Voting Beats Parallel Self-Consistency at Matched Compute
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering