Beyond the Best Guess: Improving LLM Solution Coverage with Evolution Strategies

arXiv:2608.12679 · cs.AI, cs.NE · Submitted 2026-08-13 · Read on arXiv

Conor F. Hayes, Elliot Meyerson, Kajetan Schweighofer, Roberto Dailey, Babak Hodjat, Risto Miikkulainen, Xin Qiu

Cognizant AI Lab · The University of Texas at Austin

cs.AI, cs.NE

Submitted: 2026-08-13

Updated: 2026-08-14

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: Large Language Models (LLMs) are increasingly deployed in discovery domains such as math and science.

Terminology

Summary

Large Language Models (LLMs) are increasingly deployed in discovery domains such as math and science. The usual approach is to present the problem to the model and use its answer as the proposed solution. However, beyond this best guess, discovery can be enhanced by increasing test-time compute. In a process called pass@k, the model is allowed to explore the solution space and generate diverse candidate solutions. Unfortunately, the standard approach to post-training LLMs through Reinforcement Learning (RL) may limit pass@k: the model’s output distribution narrows around high-reward outputs, causing the solution coverage to collapse. The alternative is to use Evolution Strategies (ES), a population-based, gradient-free post-training method that optimizes directly in weight space through random perturbations. As this paper shows, ES achieves consistently higher pass@k than RL and produces a broader output distribution with greater solution coverage. This coverage in turn makes it possible to achieve better results in e.g. standard math benchmarks. Thus, ES provides a better foundation for post-training in discovery problems and other domains where diverse solution coverage is critical.

The paper empirically investigates whether ES yields improved pass@k relative to RL-trained models at matched compute budgets and evaluates the implications for downstream performance through test-time scaling experiments. The main contributions are:

• The most comprehensive evaluation of ES as an LLM post-training technique to date: ES models are trained on different parameter scales from 1.5B to 32B parameters, spanning multiple model families.

• Establishing ES as a superior post-training strategy for settings requiring solution diversity: ES consistently improves pass@k over RL-trained models across all model families and scales.

• Insights into the mechanistic origins of the observed improvements in solution coverage, obtained through a systematic analysis of the output distributions of ES-trained models relative to base and RL-trained models.

• Demonstrating the value of high-quality ES output distributions compared to those of RL in several math benchmarks.

Overall, the results confirm the hypothesis that ES is better suited to pass@k than RL, and establish ES as the state-of-the-art approach to post-training in domains where diverse solution exploration is critical to performance.

ES improves pass@k over RL across model families on GSM8K. Figure 2 shows that ES consistently improves pass@k over RL across Qwen2.5-Instruct and Qwen3 models at scales ranging from 1.5B to 8B. The crossover point, where ES and RL perform equally, typically occurs at k = 2, after which ES outperforms RL. Notably, the RL pass@k curves plateau earlier than those of ES, consistent with distribution collapse narrowing the solution coverage of RL-trained models. Furthermore, for the Qwen2.5-Instruct models the base model overtakes RL at sufficiently large k, empirically confirming the findings of Yue et al. (2025). The advantage of ES over RL grows with k, suggesting that ES becomes increasingly beneficial as the test-time compute budget increases. This pattern is consistent across both model families, supporting the generality of the finding.

Crucially, the phenomenon reported by Yue et al. (2025) that the base model outperforms the RL-fine-tuned model at sufficiently large k, is not observed for ES. This suggests that ES post-training improves pass@1 without sacrificing the broad output distribution support of the base model. As a result, ES models benefit from increasing test-time compute in a way that RL models do not, making ES a potentially more suitable post-training strategy for discovery problems.

Next, pass@k for ES and RL was compared by evaluating Qwen2.5-Math-7B, Qwen2.5-14B, and Qwen2.5-32B on MATH500, Olympiad Bench, and Minerva benchmarks. Figure 3 shows that ES consistently improves pass@k over RL across Qwen2.5-Math-7B, Qwen2.5-14B, and Qwen2.5-32B. For Qwen2.5-Math-7B, ES outperforms OatZero for k > 1 across on MATH500. For Qwen2.5-14B, ES outperforms SimpleRL-Zoo for k > 1 on Olympiad Bench. Notably, for Qwen2.5-14B the base model becomes competitive with SimpleRL-Zoo at large k on Olympiad Bench. For Qwen2.5-32B, ES retains a clear advantage over SimpleRL-Zoo on Minerva where the gap grows with k. The MATH experiments provided further evidence that distribution collapse is not an artifact of a specific RL algorithm but a general consequence of RL post-training. The RL-trained models exhibit earlier pass@k saturation compared to ES, especially on Minerva, consistent with a narrowing of the output distribution support. Crucially, the base model does not outperform the ES fine-tuned model at any value of k or model scale tested, showing that ES post-training improves pass@1 without sacrificing the broad output distribution support of the base model. This pattern holds consistently from 7B to 32B parameters, indicating that ES’s ability to preserve solution coverage is not a small-scale artifact but persists as models grow substantially larger.

To better understand how ES achieves such high pass@k compared to RL or the base model, this section analyzes model behavior from three perspectives: accuracy distributions, regressions vs. progressions, and answer entropy. Figure 4 shows the accuracy distribution for ES and RL Qwen2.5-Math-7B models on MATH500 and Olympiad Bench. Each bin shows the mean accuracy across k responses for a given prompt, where the frequency on the y-axis denotes the number of prompts in each bin over the full dataset. For example, bin 1.0 contains prompts where all k responses were correct, and bin 0.0 contains prompts where all k responses were incorrect. ES and RL affect the accuracy distribution differently with respect to the base model. RL increases the frequency in bin 1.0 relative to the base model, reflecting its objective of maximizing pass@1. ES also increases frequency in bin 1.0, but the improvement is more distributed across bin (0.9, 1.0], consistent with the pass@1 differences between ES and RL reported in Section 3.2. A notable artifact of RL training is an increase in the frequency of bin 0.0 relative to the base model, which was also observed by Yue et al. (2025). This indicates that RL renders a subset of prompts that the base model could solve entirely unsolvable across k samples, reducing solution coverage and limiting the model’s ability to generalize across a broad range of problems. RL’s increase in unsolvable prompts imposes a hard ceiling on achievable pass@k performance that no amount of additional sampling can overcome. In contrast, ES reduces the frequency in bin 0.0 relative to the base model, improving solution coverage across all benchmarks.

The narrowing and broadening of solution coverage can be quantified through model progressions and regressions. Progressions measure knowledge gained during fine-tuning: prompts the base model got wrong but the fine-tuned model answered correctly. Regressions measure knowledge lost during fine-tuning: prompts the base model answered correctly but the fine-tuned model got wrong. Ideally, a post-trained model maximizes progressions and minimizes regressions. Figures 5a, 5b, and 5c report progressions and regressions for ES and RL fine-tuned Qwen2.5-Math-7B models. ES yields significant increases of progressions compared to RL: There are 6 for ES vs. 3 for RL on MATH500; 56 vs. 33 on Olympiad Bench; and 22 vs. 16 on Minerva. A natural concern when comparing ES to RL is that ES may preserve the base model distribution simply because it is not learning effectively. The progression results show ES adds more new correct solutions than RL, i.e. ES improves the model’s knowledge rather than simply not damaging it. ES also results in significantly fewer regressions compared to RL. For Qwen2.5-Math-7B, ES has 3 regressions vs. 17 for RL on MATH500; on Olympiad Bench, ES has 15 vs. 49 for RL; on Minerva, ES has 6 vs. 40 for RL. Regressions are a signal of catastrophic forgetting, where knowledge already contained within the base model is lost. RL overwrites some of this knowledge in order to maximize reward on the training distribution, whereas ES pushes parameters toward flat regions of weight space and is therefore less destructive to existing knowledge. Notably, the gains in progressions from ES are observed alongside reductions in regressions, meaning ES is not trading one for the other but improving on both simultaneously. This finding suggests ES fine-tuned models are more quality-preserving with respect to the output distribution and may generalize better to out-of-distribution problems. Together, the progression and regression results reveal that ES achieves a more favorable knowledge tradeoff than RL during fine-tuning. RL gains new knowledge at the cost of overwriting existing knowledge, while ES adds new knowledge while preserving more of what the base model already knew. In other words, ES finds regions of parameter space that are both more capable and more knowledge-preserving. Critically, this dual advantage directly explains the pass@k results, where fewer regressions broaden the lower end of the solution coverage distribution, while more progressions expand the upper end, together producing the consistently higher pass@k curves observed for ES across all benchmarks and model scales.

Figure 6 shows the entropy over answer distributions for ES and RL regressions for Qwen2.5-Math-7B. The answer distribution is constructed by taking the unique final answers per problem and computing a frequency distribution, over which the Shannon entropy is calculated. ES has a total of 24 regressions across MATH500, Olympiad Bench, and Minerva, while RL has 106. Figure 6a shows that ES maintains higher entropy than RL, indicating that when ES fails it preserves greater diversity over candidate answers. Figure 6b presents the distribution over entropy values, where each bin corresponds to the fraction of total regressions across all three benchmarks. Notably, more than 25% of RL regressions have entropy close to zero, meaning the model collapses to a small number of responses and is confidently incorrect. In contrast, the ES entropy distribution is shifted to the right relative to RL, with a higher median entropy, further confirming that ES failures are characterized by uncertainty rather than confident convergence to incorrect answers. Figure 7 provides a complementary view, presenting box plots of the entropy over the answer distribution per prompt across three subsets of problems: those that both ES and RL solve, those that only ES solves, and those that only RL solves. On problems that both methods solve, ES and RL reduce entropy relative to the base model, indicating increased confidence on solvable problems. A striking difference emerges on problems where methods fail: on problems that ES solves but RL fails, RL exhibits low entropy, indicating that the model confidently converges to a small subset of incorrect answers. In contrast, on problems that RL solves but ES fails, ES preserves or exceeds the entropy of the base model, maintaining solution diversity even in its failure modes. This asymmetry reveals a fundamental difference in how ES and RL fail: whereas RL fails confidently and narrowly, ES fails with preserved uncertainty, a property that is more amenable to correction through increased test-time compute.

Beyond its role as a tool for understanding distribution quality, pass@k can be put to good use through TTS. With high-quality output distributions, better solutions are more likely to surface over increased inference steps. However, pass@k itself is a practical TTS method only in verifiable test settings, e.g., in coding when there are a given set of tests that can be used to programatically judge whether a particular solution is correct. In more general settings, the simplest and most well-established TTS method is Self-Consistency, i.e. generating multiple independent samples and selecting an answer by plurality vote. This section evaluates ES fine-tuning in this role. The same math benchmarks are still used for this evaluation, however, instead of checking whether any of the k generated answers are correct, self-consistency is first used to identify one answer among the k generated. The evaluation thus measures how well TTS can be used with ES-generated output distributions in non-verifiable test settings. Self-Consistency was applied using the same models, parameters, and data as Section 3.2. The results mirror those in Section 4.2: As k increases, ES overtakes RL in terms of accuracy (Figures 8a, 8b, and 8c). These results indicate that not only is ES more likely to include the correct answer in its distribution, but its maximum probability answer is more likely to be correct. It is notable that on two of the three benchmarks the maximum probability answers of RL are worse than the base model, consistent with other observations of distribution degradation. Coupled with the pass@k results, these results suggest that the higher-quality distributions of ES may be generally more useful in TTS than those of RL. Exploring how ES can improve other forms of TTS, such as agentic harnesses or tree search is a promising avenue of future work.

As the importance of TTS continues to grow, methods that explicitly optimize for use in TTS settings must be developed. ES is a natural fit given its gradient-free nature, which allows a TTS objective to be defined and optimized directly. For example, ES could be used to optimize the pass@k metric during training, directly aligning the training objective with the downstream metric that matters in TTS deployments. Beyond pass@k, this framework could extend to other TTS objectives such as self-consistency voting accuracy or multi-step agentic success rate.

As LLMs are increasingly used for discovery problems, it is important that post-training methods do not degrade solution coverage. This paper shows that while RL improves pass@1 accuracy, it does so at the cost of solution coverage through distribution collapse. In contrast, ES preserves broad solution coverage while simultaneously increasing pass@1 accuracy. The benefits of ES for test-time scaling are demonstrated through pass@k performance, output distribution analysis, and downstream test-time scaling experiments, where ES consistently outperforms RL. These results position ES as a promising alternative to RL for post-training in discovery problems such as science, math, and coding, and in other settings where solution diversity is critical.

Improvements for AI systems

Improvements to AI Systems Based on This Paper:

  1. Implement Evolution Strategies (ES) as a post-training alternative to Reinforcement Learning (RL) for discovery tasks. The improved AI system will use gradient-free, population-based optimization in weight space via random perturbations, rather than reward-maximizing RL that narrows output distributions. This directly prevents distribution collapse, preserving broad solution coverage while improving pass@1 accuracy.

  2. Optimize the training objective for test-time scaling metrics like pass@k. The improved system can define and directly optimize pass@k (or other TTS objectives like self-consistency voting accuracy) during post-training, since ES is gradient-free and can handle non-differentiable objectives. This aligns training with the actual deployment metric, ensuring the model improves as more compute is spent at inference.

  3. Reduce catastrophic forgetting during fine-tuning. By pushing parameters toward flat regions of weight space, ES-trained models show significantly fewer regressions (e.g., 3 vs. 17 on MATH500) compared to RL. The improved system will preserve existing knowledge from the base model while adding new capabilities, leading to better generalization on out-of-distribution problems.

  4. Maintain high output entropy on failure cases. Unlike RL, which fails confidently and narrowly (over 25% of regressions have near-zero entropy), ES-trained models fail with preserved uncertainty and diverse candidate answers. The improved system will be more amenable to correction through additional sampling or self-consistency, as its failure modes are not collapsed to a single wrong answer.

  5. Enable better self-consistency and plurality voting performance. The improved system, when using ES post-training, will produce output distributions where the maximum-probability answer is more likely correct as k increases, outperforming RL-trained models on math benchmarks (GSM8K, MATH500, Olympiad Bench, Minerva) across scales from 1.5B to 32B parameters.

  6. Support scalable test-time compute without diminishing returns. The improved system will show pass@k curves that continue improving with larger k, rather than plateauing early like RL models. This makes it suitable for agentic harnesses, tree search, or any setting where multiple solution attempts are feasible.

  7. Provide a dual advantage in knowledge tradeoff. The improved system will simultaneously increase progressions (new correct solutions) and decrease regressions (lost solutions), unlike RL which trades one for the other. This yields higher-quality output distributions that are both more capable and more knowledge-preserving.

What the Improved AI System Can Do:

  • Solve math and science problems with higher success rates when multiple attempts are allowed, especially at large test-time compute budgets.

  • Maintain solution diversity on hard problems, avoiding the confidently wrong failure mode common in RL-trained models.

  • Generalize better to unseen problem types by preserving base model knowledge while adding new reasoning capabilities.

  • Leverage self-consistency voting more effectively, producing more reliable final answers in non-verifiable settings.

  • Scale from 1.5B to 32B parameters without losing the benefits of broad solution coverage, making it suitable for both small and large deployment scenarios.

Abstract

Large Language Models (LLMs) are increasingly deployed in discovery domains such as math and science. The usual approach is to present the problem to the model and use its answer as the proposed solution. However, beyond this best guess, discovery can be enhanced by increasing test-time compute. In a process called pass@k, the model is allowed to explore the solution space and generate diverse candidate solutions. Unfortunately, the standard approach to post-training LLMs through Reinforcement Learning (RL) may limit pass@k: the model's output distribution narrows around high-reward outputs, causing the solution coverage to collapse. The alternative is to use Evolution Strategies (ES), a population-based, gradient-free post-training method that optimizes directly in weight space through random perturbations. As this paper shows, ES achieves consistently higher pass@k than RL and produces a broader output distribution with greater solution coverage. This coverage in turn makes it possible to achieve better results in e.g. standard math benchmarks. Thus, ES provides a better foundation for post-training in discovery problems and other domains where diverse solution coverage is critical.

Sources

Related papers