Matching Accuracy, Different Geometry: Evolution Strategies vs GRPO in LLM Post-Training
summary
The gist
" Evolution Strategies (ES) serve as a scalable, gradient-free alternative to reinforcement learning methods like Group Relative Policy Optimization (GRPO) for fine-tuning Large Language Models
In short
The paper examines Evolution Strategies (ES) and GRPO in LLM post-training. ES achieves high peak accuracy but struggles with forgetting prior knowledge during sequential learning, while its updates are broad. GRPO offers targeted adjustments and stability. The hosts conclude that the choice between these two methods depends on whether an application requires peak performance or stable, continuous learning.
Key concepts
- Evolution Strategies (ES)
- ES achieves impressive peak task accuracy but lacks stability in sequential learning. It uses large, broad updates that act like a random walk through weakly informative subspaces. This behavior can lead to massive off-manifold diffusion and the forgetting of previously learned knowledge.
- GRPO
- GRPO is characterized by highly targeted adjustments. It offers greater stability than ES, making it valuable for maintaining robust systems. Its focused approach allows for high fidelity in critical areas where precision is paramount for a reliable AI system.
- Orthogonal Updates
- The authors found that the solutions derived from ES and GRPO are nearly orthogonal in direction. This means their core update vectors are almost completely unrelated, indicating they do not move toward the same target cohesively, despite sometimes achieving similar performance.
Terminology used across episodes
This episode discusses
- Matching Accuracy, Different Geometry: Evolution Strategies vs GRPO in LLM Post-Training · Paper Radio
- Evolutionary Strategies lead to Catastrophic Forgetting in LLMs
- Retaining by Doing: The Role of On-Policy Data in Mitigating Forgetting
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- ESSA: Evolutionary Strategies for Scalable Alignment · Paper Radio
- The Blessing of Dimensionality in LLM Fine-tuning: A Variance-Curvature Perspective
- Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning
- Evolution Strategies as a Scalable Alternative to Reinforcement Learning
- Evolution Strategies at the Hyperscale
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- RL's Razor: Why Online Reinforcement Learning Forgets Less
The paper
Matching Accuracy, Different Geometry: Evolution Strategies vs GRPO in LLM Post-Training · Read on arXiv
William Hoy, Binxu Wang, Xu Pan
University of Miami · Kempner Institute · Harvard University
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Matching Accuracy, Different Geometry: Evolution Strategies vs GRPO in LLM Post-Training".
Jane: The paper was written by William Hoy, Binxu Wang and Xu Pan from University of Miami and Kempner Institute and Harvard University.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Paper discussion segment 1: Tom: So, as we looked at the initial results for single-task training—training the model only on one specific task—the performance of ES is quite impressive. For instance, in Chemistry, ES achieved a very high accuracy of seventy-six point five percent.
Jane: But that story changes when we look at how these models perform in sequential learning, where they must master four different subjects like Countdown or Math one after the two. The paper shows that while ES often achieves higher peak accuracy on the new task compared to GRPO, this performance is not always maintained.
Lu: That sequential training is where the real nuance comes in for continual learning applications. The authors show that ES struggles with forgetting earlier tasks as it progresses through the pipeline, which is a major concern if we want our models to keep what they already know.
Meng: It’s definitely a trade-off there. If we use ES to achieve peak performance on the latest task, we risk losing knowledge from previous tasks in the sequential training pipeline. The results quantify this loss of general knowledge quite well.
Lalam: That suggests that for maintaining a robust and reliable AI system, the stability that GRPO provides might be more valuable than just chasing the absolute highest peak performance offered by ES.
Jane: It’s interesting how they frame this as a comparison different paths to achieving similar accuracy, right? But while the results are telling us about forgetting, we need to understand *why* these two methods behave so differently in their underlying mechanics.
Paper discussion segment 2: Tom: The core of the findings is not just that ES and GRPO are different; we have a deep geometric explanation for why they can both be correct but fundamentally distinct, which the authors call "Different Geometry." This is where the technical improvements start to shine.
Jane: The paper suggests that ES updates are much larger and broader than GRPO's. You can think of it as if ES is throwing a wide net of changes while GRPO is using a highly targeted laser to make its adjustments.
Lu: My take on this is that ES, by having these large, diffuse updates, seems to be exploring the entire parameter space in a more random way—it’s essentially acting like a random walk through weakly informative subspaces.
Meng: But if ES is behaving like a random walk and GRPO is highly targeted, how does that actually improve our ability to select the right method for this application? Do we need something that's both wide and targeted?
Lalam: It seems the authors are suggesting that we can leverage the focused, low-dimensional updates of GRPO when precision is paramount. This focus allows for high fidelity in critical areas.
Jane: And Tom mentioned this earlier—we need to look at how these large ES updates affect the model’s overall knowledge and whether that wide net approach is actually beneficial to its long-term memory.
Tom: We are going to shift our focus now toward the how this geometric difference translates into a practical recommendation for the next step, which is where we look at real improvements.
Paper discussion segment 3: Tom: The paper provides clear empirical evidence that ES can match or even exceed GRPO in task accuracy, but it also offers a way to understand *why* those solutions are so different geometrically. The authors found that these two methods are "linearly connected" in terms of performance, meaning there is no catastrophic loss barrier between them.
Jane: The fact that they are lineally connected is reassuring. It suggests we can't simply say one method is fundamentally impossible because the other isn't doing it; the path from ES to GRPO exists smoothly in terms accuracy.
Lu: But even though the performance looks smooth, we have a massive geometric separation—the authors find that ES and GRPO solutions are nearly orthogonal in direction. This means their update vectors are almost completely unrelated to each other at their core.
Meng: Orthogonal updates sound like a nightmare for practical implementation because they aren't working together cohesively; they don't move toward the same target in the same way. How do we manage two methods that disagree so fundamentally on how to change the model?
Lalam: I think this separation, coupled with ES’s random walk behavior, tells us that we need to be very intentional about which method we choose for a continuous learning environment. We must decide if exploration or precision is the primary goal.
Jane: And Tom just added that because ES accumulates these large changes in those weakly informative directions, it's great for exploring new ideas but bad for stability.
Tom: So, let's look at how this geometry allows us to choose the right tool by tying these findings to practical implications and future work.
Conclusion: Tom: We’ve seen that ES can achieve higher accuracy than GRPO on certain tasks, but we also know that the solutions found by these two methods are dramatically different in terms of size, sparsity, and direction. The paper explored the theory behind this difference—the random walk behavior of ES—and how it leads to massive off-manifold diffusion.
Jane: It’s clear that "Matching Accuracy, Different Geometry: Evolution Strategies vs GRPO in LLM Post-Training" is a crucial paper because it proves that two methods can achieve the same result without achieving anything alike in parameter space.
Lu: This separation is a huge piece of the puzzle for AI safety and reliability; since we now have data on how both methods behave, we can better predict where they might fail or succeed when they are deployed.
Meng: For us engineers, knowing that ES operates on a much larger scale and GRPO on a tightly constrained one allows us to start designing specific experimental setups tailored to the risk tolerance of our application. We can finally build systems with informed trade-offs.
Lalam: I think this paper is incredibly important for the future of AI because it shows we can achieve similar task alignment even with fundamentally different paths, which gives us more flexibility in how we guide these powerful models.
Tom: And looking at all this, I hope that when we revisit LLM post-training, "Matching Accuracy, Different Geometry: Evolution Strategies vs GRPO in LLM Post-Training" offers a better framework for understanding the trade-offs between stability and peak performance.
Lu: It’s definitely giving us something to think about regarding how our AI models learn and adapt over time.
Meng: I’m glad we can start building more targeted systems now that we understand these massive differences in update geometry.
Lalam: We are truly excited to see the culture of AI evolve with this deep understanding, thank you all for sharing your insights today on "Matching Accuracy, Different Geometry: Evolution Strategies vs GRPO in LLM Post-Training."
More episodes
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language