Evolutionary Ensemble of Agents
summary
The gist
solvers S = (s j, l j s, v j s) (sets of code files within a base repository B, with evaluation logs and scores from task evaluator f) and agents A = (a i, l i a, v i a) (agents with cumulative
In short
The episode discusses the paper "Evolutionary Ensemble of Agents" by Zongmin Yu and Liu Yang from NUS. The paper proposes an ensemble of AI coding agents that evolve together based on performance scores, using an Elo rating system to guide improvement. The hosts conclude that continuous evolution beats fixed agents, emphasizing the importance of how AI systems are organized.
Key concepts
- Evolutionary Ensemble of Agents
- This system involves a group of capable AI agents working together. These agents change their strategies and skills over time based on which ones produce better results, allowing the team to improve collectively.
- Synchronous Race
- This scoring method treats every agent as starting with the same baseline. The agent that achieves the greatest improvement over that baseline in a round wins, similar to a competition where ingredients are identical for all participants.
- Phase Mismatch
- This occurs when taking the best evolved agent from one run and using it again performs worse than the original fixed agent. This happens because the frozen agent loses the ability to perform necessary early-stage exploration required when starting a new problem.
- Integrated Agent Workspace
- This is a key design improvement where all components—solvers, scores, logs, and code repository—are merged into one stage. This allows an agent to see the consequences of its strategies directly and adjust on the fly.
Terminology used across episodes
This episode discusses
- Evolutionary Ensemble of Agents · Paper Radio
- CodeEvolve: an open source evolutionary coding agent for algorithmic discovery and optimization · Paper Radio
- VICON: Vision In-Context Operator Networks for Multi-Physics Fluid Dynamics Prediction
- AdaEvolve: Adaptive LLM Driven Zeroth-Order Optimization
- ShinkaEvolve: Towards Open-Ended And Sample-Efficient Program Evolution
- EvoX: Meta-Evolution for Automated Discovery
- Escher-Loop: Mutual Evolution by Closed-Loop Self-Referential Optimization
- AlphaEvolve: A coding agent for scientific and algorithmic discovery
- TerraLingua: Emergence and Analysis of Open-endedness in LLM Ecologies
- CORAL: Towards Autonomous Multi-Agent Evolution for Open-Ended Discovery · Paper Radio
- A Self-Improving Coding Agent
- Huxley-G"odel Machine: Human-Level Coding Agent Development by an Approximation of the Optimal Self-Improving Machine
- ThetaEvolve: Test-time Learning on Open Problems
- Group-Evolving Agents: Open-Ended Self-Improvement via Experience Sharing
- Graph In-Context Operator Networks for Generalizable Spatiotemporal Prediction
- CoEvoSkills: Self-Evolving Agent Skills via Co-Evolutionary Verification
- Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents
- Hyperagents
The paper
Evolutionary Ensemble of Agents · Read on arXiv
Zongmin Yu, Liu Yang
National University of Singapore
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Evolutionary Ensemble of Agents".
Jane: The paper was written by Zongmin Yu and Liu Yang from National University of Singapore.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back to the show, everyone. Today we're looking at a paper that's been making the rounds on arXiv, and it's called "Evolutionary Ensemble of Agents." Jane, I've got to say, the title alone got me excited — it sounds like they're trying to make AI that can improve itself, which is kind of the holy grail, right?
Jane: Absolutely, Tom. And the idea is actually pretty elegant when you break it down. You've got these really capable AI coding agents — the kind that can already write and edit code on their own — and instead of trying to build a better agent from scratch, this paper says, let's take a bunch of good ones and let them evolve together. Like a team that learns to work better over time.
Tom: So it's not about reinventing the wheel, it's about getting the wheels to coordinate?
Jane: Exactly. They call it an ensemble, which just means a group working together. And the key twist is that the agents themselves change over time — their guidance, their strategies, their skills — based on what's actually working. It's like a workplace where the employees don't just do their jobs, they also update the training manual as they learn.
Tom: And that's what makes it "evolutionary." The agents that produce better results get higher scores, and those scores influence which agents get to work on the next round of problems. So the good ideas spread, and the bad ones fade out.
Jane: Right. And the paper's authors — Zongmin Yu and Liu Yang from the National University of Singapore — they tested this on a real research problem. They were trying to fix a bottleneck in something called In-Context Operator Networks, which is a way for AI to learn mathematical operators from examples. And the ensemble actually discovered a solution that worked better than anything a single, fixed agent could find.
Tom: That's the part that gets me — it's not just a theoretical framework. They applied it to a concrete problem and it worked. So the question is, how much further can this go? If you can have a team of AI agents that keeps improving itself, what can't it do?
Jane: Well, that's what we're going to dig into. The title promises a lot, but the real story is in how they made it work — and what happens when you take the evolution away.
Summary: Tom: So Jane, we've talked about the title, but let's get into what this paper actually does. "Evolutionary Ensemble of Agents" — the summary is that they built a system where two populations evolve together: the solvers, which are actual code solutions, and the agents, which are the AI systems that write those solutions.
Jane: And the clever part is how they score the agents. It's not just about whether an agent produces a good solution in isolation. It's about whether that agent produces a better solution than another agent working on the exact same starting point. They call it a synchronous race — every agent gets the same baseline, and the one that squeezes out the most improvement wins the round.
Tom: So it's like a cooking competition where everyone gets the same ingredients, and the winner is the one who makes the best dish. That way you know the difference is the chef, not the ingredients.
Jane: Precisely. And they use an Elo rating system — you know, like chess ratings — to keep track of which agents are performing well. The agents that consistently improve on the baseline get higher ratings, and higher-rated agents get picked more often to work on the next round.
Tom: But here's the thing that really stood out to me. They ran three different versions of the experiment. One where the agents evolve continuously, one where a single fixed agent does all the work, and one where they take the best evolved agent from the first run and freeze it. And the results were pretty striking.
Jane: The evolving ensemble won. But what's more interesting is what happened with the frozen best agent. You'd think that taking the best agent from a successful run and using it again would give you a head start. But it actually performed worse — sometimes even worse than the fixed initial agent. The authors call this "phase mismatch."
Tom: Phase mismatch — that's a great way to put it. The agent that was great at the late stages of the first run was optimized for a situation where the solvers were already pretty good. But when you start a fresh search from scratch, you need early-stage strategies — exploration, trying wild ideas. The frozen agent had lost those.
Jane: And that's the core insight of the paper. It's not enough to find a good agent. You need agents that can adapt as the problem changes. The search landscape shifts — early on, you're exploring broadly; later, you're refining. A single agent, no matter how good, can't do both well.
Tom: So the takeaway is that the evolution itself is the magic, not any particular evolved agent. That's a pretty profound statement about how we should think about AI systems.
Jane: It is. And it raises a big question: if continuous adaptation is that important, how do we build systems that can do it reliably at scale? That's what we'll dig into next.
Improvements: Tom: So we've established that "Evolutionary Ensemble of Agents" shows continuous evolution beats any fixed agent. But what I want to know is, how did they actually implement this? What's the concrete improvement over what came before?
Jane: Great question, Tom. The key improvement is what they call the "integrated agent workspace." Earlier systems, like Escher-Loop, evolved code blocks and prompts in separate phases. But EvE — that's what they call their system — merges everything into one unified stage. The agent gets to see the solvers, the scores, the logs, and the base code repository all at once, and it can improve both the code and its own guidance in a single session.
Lu: If I can jump in here — that's actually a really important design choice. When you decouple solver improvement from agent improvement, the agent doesn't have full context about what's working and why. By integrating them, the agent can directly observe the consequences of its own strategies and adjust on the fly. It's like learning by doing, rather than learning by reading a report.
Meng: But I want to ask about the practical side. You're running multiple agents in parallel, each doing a full coding session, and then you're training neural networks to evaluate the results. That sounds expensive. What's the actual compute cost?
Jane: They report it pretty transparently. Each iteration takes about forty to sixty minutes, with two agents working in parallel. A full run of fifteen iterations takes about ten to fifteen hours and consumes twenty to thirty A40 GPU-hours. And the token cost — the amount of text sent to the AI — is comparable to a single large-scale training job.
Meng: Okay, that's not trivial, but it's not crazy either. For a research problem that could take a human weeks to solve, spending a day of compute is actually reasonable.
Lu: And the results justify it. They discovered a "rescale-then-interpolate" mechanism for positional encoding that let the model generalize to more examples than it was trained on. The baseline model's error jumps from about zero point zero five to zero point nine when you go beyond the training range — that's an eighteen-fold degradation. The evolved solution keeps the error below zero point one five even at the most extreme test case.
Tom: Eighteen-fold — that's not a small gap. And the interesting thing is that two independent runs of EvE both converged to the same fundamental mechanism, even though they went through different intermediate steps. That suggests the discovery is robust, not just a lucky accident.
Jane: Right. And the per-example-count curves show that EvE doesn't just improve the average — it specifically fixes the out-of-distribution regime where the baseline collapses. The other variants improve things, but they don't get the same stability.
Meng: So the improvement isn't just "agents evolve" — it's that the integration of context, the parallel evaluation, and the Elo-based credit assignment all work together to produce something that's actually usable. That's the kind of engineering that makes a difference.
Lu: And it opens the door to something bigger. If you can nest these ensembles — have an ensemble act as a single agent inside a larger ensemble — you could scale this up to much more complex problems. The paper explicitly mentions that as a future direction.
Conclusion: Tom: Well, we've covered a lot of ground on "Evolutionary Ensemble of Agents." Let's wrap it up. Jane, what's the one thing you want listeners to remember?
Jane: I think it's this: the paper shows that the way you organize AI agents matters just as much as the individual agents themselves. A team that keeps learning and adapting will beat a single brilliant but static expert. The authors proved it with a real research problem, and the results were clear — continuous evolution won every time.
Tom: And Lu, you had a bigger vision for where this goes?
Lu: Absolutely. The paper's final section talks about optimizing the connections between agents — like tuning the couplings in a physical system. If we can get that right, we might see ensembles that don't just solve math problems but actually do coherent scientific reasoning at scale. That's the frontier.
Meng: From an engineering standpoint, the fact that they published the repository and the artifacts means other teams can build on this immediately. That's how progress happens.
Tom: So we're saying goodbye to "Evolutionary Ensemble of Agents" — a paper that took a simple idea, evolution, and applied it to the way AI agents work together. It's not flashy, but it's the kind of work that could quietly change how we build AI systems.
Jane: And that's a great note to end on. Thanks for listening, everyone. We'll be back with the next paper soon. Until then, keep exploring.
Tom: Take care, folks.
More episodes
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language