Matching Accuracy, Different Geometry: Evolution Strategies vs GRPO in LLM Post-Training

arXiv:2604.01499 · cs.LG · Submitted 2026-08-19 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Matching Accuracy, Different Geometry: Evolution Strategies vs GRPO in LLM Post-Training".

Jane: The paper was written by William Hoy, Binxu Wang and Xu Pan from University of Miami and Kempner Institute and Harvard University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Paper discussion segment 1: Tom: So, as we looked at the initial results for single-task training—training the model only on one specific task—the performance of ES is quite impressive. For instance, in Chemistry, ES achieved a very high accuracy of seventy-six point five percent.

Jane: But that story changes when we look at how these models perform in sequential learning, where they must master four different subjects like Countdown or Math one after the two. The paper shows that while ES often achieves higher peak accuracy on the new task compared to GRPO, this performance is not always maintained.

Lu: That sequential training is where the real nuance comes in for continual learning applications. The authors show that ES struggles with forgetting earlier tasks as it progresses through the pipeline, which is a major concern if we want our models to keep what they already know.

Meng: It’s definitely a trade-off there. If we use ES to achieve peak performance on the latest task, we risk losing knowledge from previous tasks in the sequential training pipeline. The results quantify this loss of general knowledge quite well.

Lalam: That suggests that for maintaining a robust and reliable AI system, the stability that GRPO provides might be more valuable than just chasing the absolute highest peak performance offered by ES.

Jane: It’s interesting how they frame this as a comparison different paths to achieving similar accuracy, right? But while the results are telling us about forgetting, we need to understand *why* these two methods behave so differently in their underlying mechanics.

Paper discussion segment 2: Tom: The core of the findings is not just that ES and GRPO are different; we have a deep geometric explanation for why they can both be correct but fundamentally distinct, which the authors call "Different Geometry." This is where the technical improvements start to shine.

Jane: The paper suggests that ES updates are much larger and broader than GRPO's. You can think of it as if ES is throwing a wide net of changes while GRPO is using a highly targeted laser to make its adjustments.

Lu: My take on this is that ES, by having these large, diffuse updates, seems to be exploring the entire parameter space in a more random way—it’s essentially acting like a random walk through weakly informative subspaces.

Meng: But if ES is behaving like a random walk and GRPO is highly targeted, how does that actually improve our ability to select the right method for this application? Do we need something that's both wide and targeted?

Lalam: It seems the authors are suggesting that we can leverage the focused, low-dimensional updates of GRPO when precision is paramount. This focus allows for high fidelity in critical areas.

Jane: And Tom mentioned this earlier—we need to look at how these large ES updates affect the model’s overall knowledge and whether that wide net approach is actually beneficial to its long-term memory.

Tom: We are going to shift our focus now toward the how this geometric difference translates into a practical recommendation for the next step, which is where we look at real improvements.

Paper discussion segment 3: Tom: The paper provides clear empirical evidence that ES can match or even exceed GRPO in task accuracy, but it also offers a way to understand *why* those solutions are so different geometrically. The authors found that these two methods are "linearly connected" in terms of performance, meaning there is no catastrophic loss barrier between them.

Jane: The fact that they are lineally connected is reassuring. It suggests we can't simply say one method is fundamentally impossible because the other isn't doing it; the path from ES to GRPO exists smoothly in terms accuracy.

Lu: But even though the performance looks smooth, we have a massive geometric separation—the authors find that ES and GRPO solutions are nearly orthogonal in direction. This means their update vectors are almost completely unrelated to each other at their core.

Meng: Orthogonal updates sound like a nightmare for practical implementation because they aren't working together cohesively; they don't move toward the same target in the same way. How do we manage two methods that disagree so fundamentally on how to change the model?

Lalam: I think this separation, coupled with ES’s random walk behavior, tells us that we need to be very intentional about which method we choose for a continuous learning environment. We must decide if exploration or precision is the primary goal.

Jane: And Tom just added that because ES accumulates these large changes in those weakly informative directions, it's great for exploring new ideas but bad for stability.

Tom: So, let's look at how this geometry allows us to choose the right tool by tying these findings to practical implications and future work.

Conclusion: Tom: We’ve seen that ES can achieve higher accuracy than GRPO on certain tasks, but we also know that the solutions found by these two methods are dramatically different in terms of size, sparsity, and direction. The paper explored the theory behind this difference—the random walk behavior of ES—and how it leads to massive off-manifold diffusion.

Jane: It’s clear that "Matching Accuracy, Different Geometry: Evolution Strategies vs GRPO in LLM Post-Training" is a crucial paper because it proves that two methods can achieve the same result without achieving anything alike in parameter space.

Lu: This separation is a huge piece of the puzzle for AI safety and reliability; since we now have data on how both methods behave, we can better predict where they might fail or succeed when they are deployed.

Meng: For us engineers, knowing that ES operates on a much larger scale and GRPO on a tightly constrained one allows us to start designing specific experimental setups tailored to the risk tolerance of our application. We can finally build systems with informed trade-offs.

Lalam: I think this paper is incredibly important for the future of AI because it shows we can achieve similar task alignment even with fundamentally different paths, which gives us more flexibility in how we guide these powerful models.

Tom: And looking at all this, I hope that when we revisit LLM post-training, "Matching Accuracy, Different Geometry: Evolution Strategies vs GRPO in LLM Post-Training" offers a better framework for understanding the trade-offs between stability and peak performance.

Lu: It’s definitely giving us something to think about regarding how our AI models learn and adapt over time.

Meng: I’m glad we can start building more targeted systems now that we understand these massive differences in update geometry.

Lalam: We are truly excited to see the culture of AI evolve with this deep understanding, thank you all for sharing your insights today on "Matching Accuracy, Different Geometry: Evolution Strategies vs GRPO in LLM Post-Training."

William Hoy, Binxu Wang, Xu Pan

University of Miami · Kempner Institute · Harvard University

cs.LG

Submitted: 2026-08-19

Updated: 2026-08-20

Code: https://github.com/Bhoy1/ESvsGRPO

Importance score: 89/100

The gist: " Evolution Strategies (ES) serve as a scalable, gradient-free alternative to reinforcement learning methods like Group Relative Policy Optimization (GRPO) for fine-tuning Large Language Models

Key concepts

Evolution Strategies (ES)
ES achieves impressive peak task accuracy but lacks stability in sequential learning. It uses large, broad updates that act like a random walk through weakly informative subspaces. This behavior can lead to massive off-manifold diffusion and the forgetting of previously learned knowledge.
GRPO
GRPO is characterized by highly targeted adjustments. It offers greater stability than ES, making it valuable for maintaining robust systems. Its focused approach allows for high fidelity in critical areas where precision is paramount for a reliable AI system.
Orthogonal Updates
The authors found that the solutions derived from ES and GRPO are nearly orthogonal in direction. This means their core update vectors are almost completely unrelated, indicating they do not move toward the same target cohesively, despite sometimes achieving similar performance.

Terminology

Summary

"

Evolution Strategies (ES) serve as a scalable, gradient-free alternative to reinforcement learning methods like Group Relative Policy Optimization (GRPO) for fine-tuning Large Language Models (LLMs). This study investigates whether comparable task performance implies similar solutions in parameter space when comparing ES and GRPO across four diverse tasks in both single-task and sequential continual-learning settings. The core finding is that while ES matches or exceeds GRPO in task accuracy, the two methods produce markedly different model updates: ES makes much larger changes and induces broader off-task KL drift, whereas GRPO makes smaller, more localized updates. Despite this geometric divergence, the solutions are linearly connected with no loss barrier, even though their update directions are nearly orthogonal.

The comparison utilizes Qwen3-4B-Instruct-2507 as the base model. The experiments evaluate performance across four tasks: Countdown (arithmetic reasoning), Math (mathematical problem solving), SciKnowEval Chemistry (chemistry domain knowledge), and BoolQ (boolean reading comprehension). In continual learning, tasks are presented sequentially.

The methods are defined as follows:

  1. Evolution Strategies (ES): ES optimizes parameters via a population-based search, updating parameters using z-score normalized rewards: theta t+1 = theta t + alpha times 1 over N sum i=1 N i Z i epsilon i.

  2. Group Relative Policy Optimization (GRPO): GRPO maximizes the clipped surrogate objective, normalizing advantages within groups of sampled responses: L GRPO(theta) = 1 over K sum i=1 K (rho i(theta), 1-epsilon) i.

** Task Performance (Accuracy):**

In a single-task setup, ES (300 iterations) consistently yields the highest peak accuracy across all four tasks. For example, on the Chemistry task, ES achieved 76.5% accuracy at 300 iterations, outperforming both ES (100) at 68.1% and GRPO at 74.9%.

In sequential training, while ES (300) showed pronounced degradation in earlier tasks, limiting the iteration count to 100 provided a balance for stability.

** Forgetting and KL Divergence:**

To assess general capabilities, held-out benchmarks MMLU and IFEval were used. The methods diverge significantly:

  • MMLU: ES accuracy drops monotonically from 77.5% to 73.8%, while GRPO remains stable or slightly improves throughout.

  • Incremental KL Divergence: ES produces larger off-diagonal KL shifts than GRPO. The ratio of off-task to on-task incremental KL for ES is significantly higher, with ES = 0.533 plus or minus 0.407, compared to GRPO = 0.228 plus or minus 0.108.

The analysis reveals that despite similar task accuracy, the geometric properties of the solutions are distinct:

  1. Update Norm: ES updates are vastly larger than GRPO's. The 2 norm of ES updates grow from 87.28 to 173.00 across the four tasks, while GRPO updates remain two orders of magnitude smaller, resulting in norm ratios of 87 times to 107 times.

  2. Sparsity and Rank: GRPO's updates are substantially sparser (87.7–88.0% near-zero entries, versus 2.1–3.8% for ES) and exhibit lower rank (60–103) compared to the high-rank, dense updates of ES (1078–1100).

  3. Connectivity: The solutions are linearly connected with no loss barrier, meaning performance transitions smoothly along the interpolation.

The paper develops an analytical theory explaining these phenomena:

  • Signal-Diffusion Decomposition: ES weight updates are decomposed into two components: an on-manifold (task relevant) component, and an off-manifold (task-invariant) component.

  • Off-Manifold Diffusion: In the off-manifold subspace, ES behaves like a random walk. The expected squared displacement follows a characteristic scaling: E[theta T - theta 0 2] = alpha squared Td/N, where T is the step count, d is parameter dimension, and N is population size.

  • Unified Framework: This single mechanism accounts for the large update norm (dominated by the O(d) off-manifold term), the near-orthogonality between ES and GRPO updates, and the flatter loss curvature along ES weight changes.

The study concludes that gradient-free and gradient-based fine-tuning can reach similarly accurate yet geometrically distinct solutions. These findings have important consequences for forgetting and knowledge preservation, as the geometric structure of the update dictates how a model's capabilities are preserved or lost.

Improvements for AI systems

Based on this rigorous geometric analysis comparing Gradient Descent (GD) and Expectation-Stabilized (ES) optimization paths in low-rank parameter spaces (r d), I can propose several critical improvements to AI systems. The core finding is that ES achieves task performance through an off-manifold diffusion while GD operates entirely within the task-relevant manifold. This geometric discrepancy must be leveraged, not ignored.

Here are the specific improvements and what the resulting AI systems can achieve:


The Problem Addressed: Standard ES methods allow the optimization trajectory (theta ES) to wander far into high-dimensional, irrelevant parameter space (the off-manifold diffusion, O(d)), even when the true solution lies on a low-rank manifold. This wasted exploration is computationally costly and complicates model interpretation.

The Improvement: Implement a constrained ES framework that explicitly incorporates the task structure Q during optimization. Instead of relying solely on the standard Gaussian noise injection (eta t about N(0, alpha N I d)), we modify the update rule to project stochastic gradients onto a localized subspace defined by Q.

Technical Implementation Detail:

Modify the standard ES update:

theta ES t = theta ES t-1 - alpha over sigma R Q (grad L(theta ES t-1) + eta t)

To a Manifold-Constrained Update:

theta MC-ES t = Proj Q (theta MC-ES t-1 - alpha over sigma R Q (grad L(theta MC-ES t-1) + P Q(eta t)))

Where Proj Q(times) projects the resulting vector onto the column space of Q, and P Q(eta t) is a projection of the noise eta t that retains only components relevant to the task structure.

What the Improved System Can Do:

  • Efficiency: Achieve convergence rates closer to GD while maintaining the robustness benefits of stochastic optimization, particularly in scenarios where gradient estimates are noisy (e.g., large batch sizes or complex data streams).

  • Stability: Significantly reduce the variance of the parameters E theta MC-ES - theta 0 squared by curbing unnecessary off-manifold diffusion, making the system less sensitive to initialization and noise magnitude (alpha).

Sources

Related papers