Evolution or Illusion? Rethinking Evaluation in LLM Evolutionary Search
cs.CL, cs.AI, cs.LG
Submitted: 2026-09-17
Updated: 2026-09-17
Code: https://github.com/codelion/openevolve
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Effective Harness Engineering for Algorithm Discovery with Coding Agents
- Efficient Prediction of Pass@k Scaling in Large Language Models
- Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
- Evolution of Heuristics: Towards Efficient Automatic Algorithm Design Using Large Language Model
- AdaEvolve: Adaptive LLM Driven Zeroth-Order Optimization
- EvoX: Meta-Evolution for Automated Discovery
- Barbarians at the Gate: How AI is Upending Systems Research
- Don't Gamble, GAMBLe: An Analytical Framework for AI-Driven Research Systems
- Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution
- Simple Baselines are Competitive with Code Evolution
- How Do Large Language Monkeys Get Their Power (Laws)?
- Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering