Evolutionary Ensemble of Agents

arXiv:2605.09018 · cs.NE, cs.AI, cs.LG · Submitted 2026-08-17 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Evolutionary Ensemble of Agents".

Jane: The paper was written by Zongmin Yu and Liu Yang from National University of Singapore.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back to the show, everyone. Today we're looking at a paper that's been making the rounds on arXiv, and it's called "Evolutionary Ensemble of Agents." Jane, I've got to say, the title alone got me excited — it sounds like they're trying to make AI that can improve itself, which is kind of the holy grail, right?

Jane: Absolutely, Tom. And the idea is actually pretty elegant when you break it down. You've got these really capable AI coding agents — the kind that can already write and edit code on their own — and instead of trying to build a better agent from scratch, this paper says, let's take a bunch of good ones and let them evolve together. Like a team that learns to work better over time.

Tom: So it's not about reinventing the wheel, it's about getting the wheels to coordinate?

Jane: Exactly. They call it an ensemble, which just means a group working together. And the key twist is that the agents themselves change over time — their guidance, their strategies, their skills — based on what's actually working. It's like a workplace where the employees don't just do their jobs, they also update the training manual as they learn.

Tom: And that's what makes it "evolutionary." The agents that produce better results get higher scores, and those scores influence which agents get to work on the next round of problems. So the good ideas spread, and the bad ones fade out.

Jane: Right. And the paper's authors — Zongmin Yu and Liu Yang from the National University of Singapore — they tested this on a real research problem. They were trying to fix a bottleneck in something called In-Context Operator Networks, which is a way for AI to learn mathematical operators from examples. And the ensemble actually discovered a solution that worked better than anything a single, fixed agent could find.

Tom: That's the part that gets me — it's not just a theoretical framework. They applied it to a concrete problem and it worked. So the question is, how much further can this go? If you can have a team of AI agents that keeps improving itself, what can't it do?

Jane: Well, that's what we're going to dig into. The title promises a lot, but the real story is in how they made it work — and what happens when you take the evolution away.

Summary: Tom: So Jane, we've talked about the title, but let's get into what this paper actually does. "Evolutionary Ensemble of Agents" — the summary is that they built a system where two populations evolve together: the solvers, which are actual code solutions, and the agents, which are the AI systems that write those solutions.

Jane: And the clever part is how they score the agents. It's not just about whether an agent produces a good solution in isolation. It's about whether that agent produces a better solution than another agent working on the exact same starting point. They call it a synchronous race — every agent gets the same baseline, and the one that squeezes out the most improvement wins the round.

Tom: So it's like a cooking competition where everyone gets the same ingredients, and the winner is the one who makes the best dish. That way you know the difference is the chef, not the ingredients.

Jane: Precisely. And they use an Elo rating system — you know, like chess ratings — to keep track of which agents are performing well. The agents that consistently improve on the baseline get higher ratings, and higher-rated agents get picked more often to work on the next round.

Tom: But here's the thing that really stood out to me. They ran three different versions of the experiment. One where the agents evolve continuously, one where a single fixed agent does all the work, and one where they take the best evolved agent from the first run and freeze it. And the results were pretty striking.

Jane: The evolving ensemble won. But what's more interesting is what happened with the frozen best agent. You'd think that taking the best agent from a successful run and using it again would give you a head start. But it actually performed worse — sometimes even worse than the fixed initial agent. The authors call this "phase mismatch."

Tom: Phase mismatch — that's a great way to put it. The agent that was great at the late stages of the first run was optimized for a situation where the solvers were already pretty good. But when you start a fresh search from scratch, you need early-stage strategies — exploration, trying wild ideas. The frozen agent had lost those.

Jane: And that's the core insight of the paper. It's not enough to find a good agent. You need agents that can adapt as the problem changes. The search landscape shifts — early on, you're exploring broadly; later, you're refining. A single agent, no matter how good, can't do both well.

Tom: So the takeaway is that the evolution itself is the magic, not any particular evolved agent. That's a pretty profound statement about how we should think about AI systems.

Jane: It is. And it raises a big question: if continuous adaptation is that important, how do we build systems that can do it reliably at scale? That's what we'll dig into next.

Improvements: Tom: So we've established that "Evolutionary Ensemble of Agents" shows continuous evolution beats any fixed agent. But what I want to know is, how did they actually implement this? What's the concrete improvement over what came before?

Jane: Great question, Tom. The key improvement is what they call the "integrated agent workspace." Earlier systems, like Escher-Loop, evolved code blocks and prompts in separate phases. But EvE — that's what they call their system — merges everything into one unified stage. The agent gets to see the solvers, the scores, the logs, and the base code repository all at once, and it can improve both the code and its own guidance in a single session.

Lu: If I can jump in here — that's actually a really important design choice. When you decouple solver improvement from agent improvement, the agent doesn't have full context about what's working and why. By integrating them, the agent can directly observe the consequences of its own strategies and adjust on the fly. It's like learning by doing, rather than learning by reading a report.

Meng: But I want to ask about the practical side. You're running multiple agents in parallel, each doing a full coding session, and then you're training neural networks to evaluate the results. That sounds expensive. What's the actual compute cost?

Jane: They report it pretty transparently. Each iteration takes about forty to sixty minutes, with two agents working in parallel. A full run of fifteen iterations takes about ten to fifteen hours and consumes twenty to thirty A40 GPU-hours. And the token cost — the amount of text sent to the AI — is comparable to a single large-scale training job.

Meng: Okay, that's not trivial, but it's not crazy either. For a research problem that could take a human weeks to solve, spending a day of compute is actually reasonable.

Lu: And the results justify it. They discovered a "rescale-then-interpolate" mechanism for positional encoding that let the model generalize to more examples than it was trained on. The baseline model's error jumps from about zero point zero five to zero point nine when you go beyond the training range — that's an eighteen-fold degradation. The evolved solution keeps the error below zero point one five even at the most extreme test case.

Tom: Eighteen-fold — that's not a small gap. And the interesting thing is that two independent runs of EvE both converged to the same fundamental mechanism, even though they went through different intermediate steps. That suggests the discovery is robust, not just a lucky accident.

Jane: Right. And the per-example-count curves show that EvE doesn't just improve the average — it specifically fixes the out-of-distribution regime where the baseline collapses. The other variants improve things, but they don't get the same stability.

Meng: So the improvement isn't just "agents evolve" — it's that the integration of context, the parallel evaluation, and the Elo-based credit assignment all work together to produce something that's actually usable. That's the kind of engineering that makes a difference.

Lu: And it opens the door to something bigger. If you can nest these ensembles — have an ensemble act as a single agent inside a larger ensemble — you could scale this up to much more complex problems. The paper explicitly mentions that as a future direction.

Conclusion: Tom: Well, we've covered a lot of ground on "Evolutionary Ensemble of Agents." Let's wrap it up. Jane, what's the one thing you want listeners to remember?

Jane: I think it's this: the paper shows that the way you organize AI agents matters just as much as the individual agents themselves. A team that keeps learning and adapting will beat a single brilliant but static expert. The authors proved it with a real research problem, and the results were clear — continuous evolution won every time.

Tom: And Lu, you had a bigger vision for where this goes?

Lu: Absolutely. The paper's final section talks about optimizing the connections between agents — like tuning the couplings in a physical system. If we can get that right, we might see ensembles that don't just solve math problems but actually do coherent scientific reasoning at scale. That's the frontier.

Meng: From an engineering standpoint, the fact that they published the repository and the artifacts means other teams can build on this immediately. That's how progress happens.

Tom: So we're saying goodbye to "Evolutionary Ensemble of Agents" — a paper that took a simple idea, evolution, and applied it to the way AI agents work together. It's not flashy, but it's the kind of work that could quietly change how we build AI systems.

Jane: And that's a great note to end on. Thanks for listening, everyone. We'll be back with the next paper soon. Until then, keep exploring.

Tom: Take care, folks.

Zongmin Yu, Liu Yang

National University of Singapore

cs.NE, cs.AI, cs.LG

Submitted: 2026-08-17

Updated: 2026-08-18

Code: https://github.com/scaling-group/eve

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 94/100

The gist: solvers S = (s j, l j s, v j s) (sets of code files within a base repository B, with evaluation logs and scores from task evaluator f) and agents A = (a i, l i a, v i a) (agents with cumulative

Key concepts

Evolutionary Ensemble of Agents
This system involves a group of capable AI agents working together. These agents change their strategies and skills over time based on which ones produce better results, allowing the team to improve collectively.
Synchronous Race
This scoring method treats every agent as starting with the same baseline. The agent that achieves the greatest improvement over that baseline in a round wins, similar to a competition where ingredients are identical for all participants.
Phase Mismatch
This occurs when taking the best evolved agent from one run and using it again performs worse than the original fixed agent. This happens because the frozen agent loses the ability to perform necessary early-stage exploration required when starting a new problem.
Integrated Agent Workspace
This is a key design improvement where all components—solvers, scores, logs, and code repository—are merged into one stage. This allows an agent to see the consequences of its strategies directly and adjust on the fly.

Terminology

Summary

Summary

The paper introduces Evolutionary Ensemble (EvE), a decentralized framework that organizes existing, highly capable coding agents into a live, co-evolving system for algorithmic discovery. Rather than reinventing the wheel within the “LLMs as optimizers” paradigm, EvE fixes the base agent substrate and focuses entirely on evolving the cumulative guidance and skills that dictate agent behaviors.

Method. EvE maintains two scored populations: solvers S = (s j, l j s, v j s) (sets of code files within a base repository B, with evaluation logs and scores from task evaluator f) and agents A = (a i, l i a, v i a) (agents with cumulative working logs and scores). In each iteration n, EvE samples a set of high-performing working agents A n, along with reference sets of solvers n and agents n, combined with the base code repository B. Each working agent a i in A n operates on the same reference set to generate new solvers and agents: (i, i, i a) = a i(n, n, B). By forcing all sampled agents to refine the same reference set, EvE constructs a strictly pairwise competition where variance in solver quality is directly attributed to each agent’s strategy. A win-loss matrix W is built from evaluation results and processed via the EloUpdate function to adjust agent scores. Following evaluation, the agent ensemble is expanded by integrating modified agents i not equal to a i along with their session logs. EvE integrates solver improvement and self-referential agent optimization into a single, unified stage, granting the working agent full visibility into examples of solvers and agents, their scores and logs, and the base code repository all at once.

Task. The paper applies EvE to a research bottleneck in In-Context Operator Networks (ICON), specifically the positional-encoding (PE) design for example-count generalization. ICON is a transformer-based model that receives k input-output example pairs at inference time and infers the hidden operator on the fly. The core bottleneck: the model is trained with a fixed number of examples (e.g., k = 5) but must generalize to larger k at inference. The original design uses a learned embedding table indexed by example slots with fixed capacity, lacking entries beyond the fifth example, causing prediction performance to degrade sharply. The benchmark is 1D conservation law with random cubic flux: d t u + d x (au cubed + bu squared + cu) = 0, with a, b, c about Uniform[-1, 1]. The headline metric averages mean absolute error over example counts k = 1 through k = 10: e = 1 over 10 sum k=1 10 e k. The solver score is s = -e. Each candidate is trained for 2,000 steps during search.

Results. Three experimental conditions were designed: EvE (full ensemble), Static-Initial (initial agent used throughout, no agent evolution), and Static-Final (single best-rated agent from a completed EvE run extracted and frozen). Each condition ran twice independently with T = 15 iterations, I = 2 working agents in parallel, J = 8 reference solvers, and K = 4 reference agents per iteration. The two EvE runs converge to almost identical final errors. The two Static-Initial runs diverge: one approaches EvE, the other plateaus higher. One Static-Final run plateaus early above even the worse Static-Initial run, while the other lands between the two Static-Initial curves. The paper attributes this to phase mismatch: the frozen agent was optimized for the late stage of the original EvE run but a fresh search requires early-stage exploration strategies. For per-example-count generalization, the Seed (vanilla ICON PE) collapses catastrophically beyond k = 5. All evolved methods avoid this collapse. EvE performs best, with error staying below 0.15 at k = 10 at 2,000 steps and below 0.08 at k = 10 under full training (10,000 steps). Both Static-Initial and Static-Final underperform EvE. The paper shows stage-dependent agent adaptation: early iterations identify structural bottlenecks and propose broad exploration; middle phases ground updates in accumulated solver evidence and promote promising PE families; late phases retract earlier strategies that stopped producing gains and redirect search toward finer-grained refinements. The two independent EvE runs exhibit the same progression, suggesting this pattern is robust. All six evolved solvers exploit the role/example decomposition of the flat token-position index p into a within-example role r = p 3 and an example index m = p/3. EvE autonomously discovered a robust rescale-then-interpolate mechanism. For computational cost, all coding-agent sessions used GPT 5.4 at medium reasoning effort via Codex. Each iteration completes in roughly 40 to 60 minutes end-to-end; over 15 iterations, one complete run takes approximately 10 to 15 hours, consuming 20-30 A40 GPU-hours.

Conclusion. The paper concludes that organizing agents into a live, self-revising ensemble is the fundamental driver for breaking through static performance ceilings. EvE’s role-free design achieves universal compatibility and naturally supports recursive nesting, allowing any existing multi-agent system or an entire ensemble to be encapsulated as a single individual within the evolutionary loop. The paper identifies a future frontier: optimization of inter-agent connection topology, analogous to an Ising model, where the goal is to achieve a phase transition into coherent large-scale scientific reasoning rather than stochastic noise or collapsed perfect alignment.

Improvements for AI systems

Based on the paper, here are the specific improvements I can implement in AI systems:


Improvement: Replace single-agent or static-agent orchestration with a decentralized, co-evolving population of coding agents. Each agent carries cumulative guidance logs and an Elo rating updated through synchronous pairwise races.

What the improved system can do:

  • Dynamically sample 2+ working agents per iteration to refine the same solver baseline, then update agent ratings based on marginal solver improvement (Elo K=32, starting at 1500).

  • Automatically promote agents that extract the most gains at the current search stage, and demote those that stall.

  • Maintain a reference set of 8 solvers and 4 agents per iteration to provide context, enabling agents to learn from both successes and failures of peers.

Summary of Capabilities: The improved system autonomously discovers novel algorithmic solutions (like the rescale-then-interpolate PE) on complex research codebases, adapts its own search strategies to the evolving problem landscape, and breaks through static performance ceilings—all within a token budget comparable to a single large training run.

Sources

Related papers