On the Effect of Sampling Diversity in Scaling LLM Inference

arXiv:2502.11027 · cs.LG · Submitted 2026-08-10 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "On the Effect of Sampling Diversity in Scaling LLM Inference".

Jane: The paper was written by Tianchun Wang, Yuanzhou Chen, Zichuan Liu, Jonathan Light, Weiyang Liu et al. from The Pennsylvania State University and University of California, Los Angeles and Carnegie Mellon University and Rensselaer Polytechnic Institute and The Chinese University of Hong Kong and Max Planck Institute for Intelligent Systems and NEC Laboratories America.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title: Tom: Welcome back, everyone. We are looking at a paper that's been making the rounds on arXiv, and it's called "On the Effect of Sampling Diversity in Scaling LLM Inference." Jane, I have to say, just reading that title gets me excited, because it's tackling something we all kind of feel but rarely see studied this directly.

Jane: Absolutely, Tom. And the title is actually perfect because it tells you exactly what the paper is about. When we ask a large language model to solve a problem, we don't just ask it once. We ask it many times and pick the best answer. That's the "scaling inference" part. The "sampling diversity" part is about how different those answers are from each other. If you ask the same question a hundred times, do you get a hundred different approaches, or do you get the same answer ninety-nine times with tiny variations?

Tom: Right, and that's the crux of it. The authors are from a bunch of places, Penn State, UCLA, Carnegie Mellon, NEC Labs. And they're basically saying that if your hundred answers are all clustered together, you're not actually exploring the solution space. You're just poking the same spot over and over.

Jane: Exactly. Think of it like a treasure hunt. If you have a hundred people looking for treasure, and they all follow the exact same map, you're probably going to find the same thing, or maybe nothing. But if you give them slightly different maps, or different starting points, you're much more likely to actually find the treasure. That's the intuition here.

Tom: And the title also hints at the fact that this isn't just an empirical trick. They're trying to build a theory around it. They want to explain *why* diversity helps, not just show that it does.

Jane: And that's what I love about this. It's not just a bunch of experiments. They have a theorem that shows, under some reasonable assumptions, that if you make your sampling more diverse, you get better performance with the same number of attempts. The math is in the appendix, but the idea is clear.

Tom: So for our listeners, the big picture is this. We've been scaling up these models by making them bigger, but there's another lever. We can scale up how we *use* them, and this paper says the quality of that usage depends heavily on how much variety we introduce.

Jane: And that's a huge deal because it means we might be able to get more out of the models we already have. We don't necessarily need a bigger model. We just need to be smarter about how we sample from the one we've got.

Tom: So, we've got the title, we've got the core idea. But the paper goes much deeper. Next up, we're going to talk about the actual summary and the main claims they're making. Stick around.

Summary: Tom: We're back with "On the Effect of Sampling Diversity in Scaling LLM Inference." Jane, we talked about the title, but the abstract of this paper is dense with ideas. What's the one-sentence version of what they're claiming?

Jane: The one-sentence version is that if you want to get the best answer out of a language model by sampling multiple times, you should actively try to make those samples different from each other, not just crank up the randomness.

Tom: And they don't just say it. They show it. They have this concept of "Best-of-N" sampling. You generate N answers, and you pick the best one. Their theorem basically says that if you diversify the prompts you use to generate those N answers, your error rate drops faster than if you just use the same prompt over and over.

Jane: Right. And they introduce this idea of a "diversity-fidelity trade-off." It's a really elegant way of thinking about it. You want your prompts to be different enough to explore new territory, but you don't want them to be so different that they're off-topic or nonsensical.

Tom: So it's a sweet spot. If you perturb the prompt too little, you get no diversity. If you perturb it too much, you get garbage. But in the middle, you get this nice boost in performance.

Jane: Exactly. And they actually test this. They have this experiment where they generate solution ideas with varying levels of relevance to the question. Irrelevant ideas, like baking tips for a math problem, don't help. Verbatim repetition of the question doesn't help. But moderately relevant ideas, like "try using the Pythagorean theorem," give a real boost.

Tom: And that's a really practical insight. It tells you how to actually design these perturbations. You can't just throw random text at the model. You have to be thoughtful about it.

Jane: The other big claim in the summary is about when this works and when it doesn't. They show it works across different temperatures, with chain-of-thought prompting, and with different verifiers. But they also found a failure mode. Majority voting, where you pick the most common answer, doesn't benefit from diversity in the same way.

Tom: That's a fascinating result. It makes sense, though. Majority voting is about consensus. If you diversify the answers, you might break the consensus. Best-of-N is about finding the one gem in the pile, so diversity helps you find that gem.

Jane: Precisely. And that's why this paper is so valuable. It's not just a "diversity is good" paper. It's a "here's when diversity is good, and here's when it's not" paper. That's the kind of nuance we need.

Tom: So we have the theory, we have the trade-off, we have the failure mode. But how do they actually implement this? What does a "diversified prompt" look like in practice? That's what we're going to dig into next.

Improvements: Tom: We're back with "On the Effect of Sampling Diversity in Scaling LLM Inference." Jane, we've talked about the theory and the trade-off. But the paper also proposes concrete ways to actually introduce this diversity. What are they?

Jane: They break it down into two main categories. The first is "task-level" perturbations. These are changes to the prompt that are the same for every question. For example, you can inject a role, like "You are a meticulous software engineer," or you can inject a strategy instruction, like "Write your code in a modular way."

Tom: So you're changing the persona or the approach, but the question itself stays the same. That's a simple, cheap way to get some diversity.

Jane: Exactly. And the second category is "query-level" perturbations. These are changes that are specific to the question being asked. The most interesting one is called "Random Idea Injection." You use a separate model, a "thinker," to generate a few solution ideas for the question. Then you inject those ideas into the prompt before asking the main model to solve it.

Tom: So instead of just saying "solve this," you're saying "here are some hints, now solve it." And each hint is different, so you get different solutions.

Jane: Right. And they have a few variants. The "Single" variant uses the same model to generate the ideas and the solutions. The "Dual" variant uses a different, potentially stronger model to generate the ideas. And the "Diverse" variant uses a whole pool of models to generate a set of ideas, and then randomly picks one for each attempt.

Tom: And the results are pretty striking. On MMLU-Pro, they got a ten point eight percent improvement over direct sampling. On MATH, eight point two percent. On HumanEval, four point seven percent. Those are significant gains just by changing the prompt.

Jane: And they also have a "Random Query Rephraser" which just restates the question in different words. That also helps, but the idea injection seems to be the more powerful technique.

Tom: So the improvements aren't just theoretical. They're practical, and they're measurable. But I'm curious about the limits. They mentioned that majority voting is a failure mode. Are there other conditions where this doesn't work?

Jane: They found that the strength of the thinker model matters. If you use a weak model to generate ideas, you get less of a boost. And the number of ideas matters too. More ideas, up to a point, leads to better performance.

Tom: So it's not a magic bullet. You need to choose your thinker wisely, and you need to generate enough ideas. But when you do it right, the gains are real.

Jane: And that's the key takeaway from the improvements section. It's a toolbox. You have task-level and query-level tools, and you need to pick the right one for the job.

Tom: So we've got the tools. But the paper also goes into a lot of detail about the conditions under which these tools work best. Let's talk about the first page of the paper and what it sets up.

First Page: Tom: We're back with "On the Effect of Sampling Diversity in Scaling LLM Inference." Jane, the first page of this paper is a masterclass in motivation. They start with this observation about non-determinism in LLMs.

Jane: Right. And it's a really interesting framing. For a long time, people saw the randomness in LLM outputs as a bug. They wanted deterministic, reproducible results. But this paper flips that on its head. They say, look, this non-determinism can be a feature, especially when you're doing test-time scaling.

Tom: And they have this great figure, Figure one that shows the difference between direct sampling and diversified sampling. Direct sampling gives you a cluster of solutions all bunched together. Diversified sampling spreads them out across the solution space.

Jane: That visual is so helpful. It makes the whole problem clear in one image. You can see that the diversified samples are covering more ground, which means they're more likely to hit the correct answer.

Tom: They also have this table, Table one that shows the effect of different injection strategies. They compare no perturbation, role injection, instruction injection, and a nonsense text called "Jabberwocky." And the results are telling.

Jane: The Jabberwocky, which is just a nonsense poem, actually hurts performance. It makes the solutions more similar to each other, and the pass rate drops. But the role and instruction injections increase diversity and improve the pass rate.

Tom: So that's the empirical motivation. It's not just theory. They show that meaningful perturbations help, and meaningless ones don't. That sets the stage for the whole paper.

Jane: And it also introduces the core puzzle. Why does diversity help? The rest of the paper is dedicated to answering that question with math and more experiments.

Tom: So the first page is really about establishing the problem and the intuition. It's saying, "Hey, look at this interesting phenomenon. Let's study it." And they do.

Jane: And they do it in a very thorough way. They don't just show that it works. They show why it works, when it works, and when it doesn't. That's the mark of a good paper.

Tom: So we've covered the title, the summary, the improvements, and the first page. Now it's time to wrap this up and give our final thoughts.

Conclusion: Tom: Alright, Jane, we've spent a good amount of time with "On the Effect of Sampling Diversity in Scaling LLM Inference." Let's bring it all together. What's the big picture here?

Jane: The big picture is that we have a new lever for improving LLM performance. We don't have to just make the models bigger. We can make our sampling strategy smarter. By introducing meaningful diversity into the prompts, we can get better answers from the same model, with the same compute budget.

Tom: And the paper gives us a framework for doing that. The diversity-fidelity trade-off is a really useful principle. It tells us that we need to be thoughtful about how we perturb prompts. Too little and we get nothing. Too much and we break the model.

Jane: And they also give us concrete tools. The task-level perturbations are cheap and easy. The query-level perturbations are more powerful but require a bit more setup. And they show us that the choice of thinker model and the number of perturbations matter.

Tom: But they also warn us. Diversity isn't a universal good. Majority voting is a place where it can actually hurt. So we need to match our sampling strategy to our verification strategy.

Jane: That's a really important practical insight. If you're using Best-of-N, diversify. If you're using majority voting, be careful. It's a nuanced message, and I appreciate that.

Tom: For me, the most exciting part is the theoretical foundation. They didn't just show that diversity helps. They proved it. That gives us confidence that this isn't a fluke. It's a fundamental property of how these models work.

Jane: And that opens up a lot of future work. We can start designing better perturbation strategies, better thinkers, and better ways to balance diversity and fidelity. This paper is a foundation, not a final answer.

Tom: So, as we say goodbye to this paper, I think the message is clear. When you're scaling inference, don't just sample more. Sample differently. That's the key to unlocking better performance.

Jane: And with that, we're ready to move on to the next paper. Thanks for listening, everyone. We'll see you next time.

Tianchun Wang, Yuanzhou Chen, Zichuan Liu, Jonathan Light, Weiyang Liu, Haifeng Chen, Xiang Zhang, Wei Cheng

The Pennsylvania State University · University of California, Los Angeles · Carnegie Mellon University · Rensselaer Polytechnic Institute · The Chinese University of Hong Kong · Max Planck Institute for Intelligent Systems · NEC Laboratories America

cs.LG

Submitted: 2026-08-10

Comments: 31 pages

Code: https://github.com/xztcwang/samplediv

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 63/100

Key concepts

Sampling Diversity
This refers to the variation between multiple answers when a language model is queried repeatedly. If the model produces clustered or similar outputs, diversity is low. The paper argues that increasing this variety allows the model to explore more of the solution space.
Diversity-Fidelity Trade-off
This concept describes finding an optimal balance when modifying prompts. Prompts must be sufficiently different to encourage exploration (diversity), but they cannot be so different that they become nonsensical or irrelevant to the original question (fidelity).
Perturbation Strategies
These are methods used to introduce intentional variation into prompts. They can be task-level changes, such as assigning a specific persona, or query-level changes, like injecting external solution ideas generated by a separate model.
Best-of-N Sampling
This is a technique where multiple answers (N) are generated from different samples. The best result is then selected. The paper demonstrates that diversifying the prompts used for these N attempts significantly reduces the error rate.

Terminology

Summary

Summary

This paper systematically investigates the effect of sampling diversity on scaling large language model (LLM) inference, particularly within the Best-of-N sampling framework. The authors provide a theoretical justification for why diversified sampling improves performance, derive a design principle for introducing diversity, instantiate several perturbation strategies, analyze the conditions under which diversified exploration remains effective, and empirically evaluate its impact across reasoning, mathematics, and code-generation tasks.

The paper begins by motivating the study with the observation that direct sampling from an LLM using the same prompt often leads to similar outputs, trapped within a local cluster, whereas diverse candidate solutions should span multiple clusters, breaking out of local clusters. The authors note that "the concentrated nature of the trapped solutions might stem from the limited diversity imposed by post-training objectives, which are typically designed to optimize zero-shot performance and align LLM as instruction-following chatbots." They present preliminary empirical evidence (Table 1) showing that perturbation strategies that reduce solution similarity (e.g., Role and Instruction injections) improve Pass@k rates on the MBPP benchmark.

The core theoretical contribution is Theorem 3.5, which states that under two hypotheses—dispersion (Hypothesis 3.1) and fidelity (Hypothesis 3.3)—introducing auxiliary diversity improves Best-of-N performance. The theorem shows that diversified sampling achieves both a lower asymptotic error rate and faster convergence as N grows. Specifically, the error reduction factor is characterized by a sequence C N = (mu 1 squared N/(1+ epsilon)), where mu 1 measures the dispersion induced by the diversity source and epsilon measures the fidelity loss. The authors explain: "This theorem implies two distinct advantages of introducing auxiliary diversity: (i) Lower asymptote. Diversity shrinks the 'blind-spot' set of instances that remain unsolved as N grows. (ii) Faster convergence. The error reduction factor improves with richer (but faithful) diversity, yielding steeper Best-of-N gains."

Building on this theorem, the authors derive a diversity-fidelity trade-off principle, which states that effective perturbations should be different enough to create exploration (boost mu 1), but faithful enough to avoid harming single-attempt quality (control epsilon). They empirically validate this principle by analyzing the relationship between perturbation-question similarity and task performance (Figure 2). The results show a non-monotonic relationship: "irrelevant ideas (Perturbation 1) and verbatim repetition (Perturbation 5) fail to improve performance and may even degrade it, while performance increases with higher relevance, peaks with task-aligned ideas (Perturbation 3), and then declines again when similarity becomes excessive."

Guided by this principle, the authors instantiate two categories of perturbations: (1) Task-level perturbations, which are task-dependent but independent of specific questions, including Role injection (e.g., mentor, optimizer, innovator) and Strategical Instruction injection (e.g., Write the code in a highly modular way...); and (2) Query-level perturbations, which are directly tied to the questions, including Random Idea Injection (RandIdeaInj), where an LLM acts as a thinker to propose task-related ideas, and Random Query Rephraser (RandQReph), which restates the input question. Both query-level strategies support three variants: Single (the model itself generates ideas), Dual (a separate model is used), and Diverse (a pool of models provides varied perturbations).

The paper then analyzes when diversified perturbations are effective across several conditions. First, regarding varying sampling temperatures, the authors find that perturbations and direct sampling all exhibit improvements at higher temperatures, and diversified sampling provides additional gains on top of temperature-induced improvements (Figure 3). Second, regarding varying thinker models, they show that stronger thinker models, such as DeepSeek-V3, raise the scaling curve (Figure 4). Third, regarding perturbation cardinality, they find that increasing the number of injection ideas raises the scaling curve, whereas using only a single injection candidate yields a noticeably lower curve (Figure 5). Fourth, under Chain-of-Thought (CoT) reasoning, they show that task-level and query-level perturbations improve performance under CoT, yielding up to a 4.7% relative gain in Pass@100 on HumanEval with GPT-4o-mini and a 7.4% relative gain on APPS with Claude-4-Sonnet (Figure 6).

The paper also examines the effect of verification methods. Under an LLM-as-a-Judge verifier, perturbations remain effective (Figure 7). However, the authors identify a failure mode: unlike pass@k, majority voting does not enjoy the same performance boost from repeated sampling and can even degrade performance in the worst case. They provide Proposition 5.1, which states that "Given one-shot correct probability p(y*) where y* is the correct answer in response set Y, as N to infinity, majority vote accuracy converges to 1 iff p(y*) = y in Y p(y), while Best-of-N accuracy converges to 1 iff p(y*) > 0. This is empirically confirmed in Figure 8, where perturbations do not yield consistent improvements over direct sampling" under majority voting on MATH.

Finally, the paper systematically evaluates the effectiveness of diversified perturbations across six benchmarks: MMLU-Pro (reasoning), GSM-Hard and MATH (mathematics), and HumanEval, MBPP, and APPS (code generation). For RandIdeaInj, the results show a 10.8% increase in reasoning on MMLU-Pro, a 8.2% increase in the mathematics on MATH, and a 4.7% increase in coding on the Humaneval dataset, over the direct sampling (Figure 9). The authors also report that higher solution diversity is accompanied by better performance, corroborating their theoretical results (Table 2). Additional results show that task-level perturbations yield notable increases of 6.7% EM@100 on MMLU-Pro, 9.6% EM@100 on MATH, and 9.5% Pass@100 on MBPP (Figure 10), and that combining perturbations (e.g., Instruction + Dual) further enhances performance across multiple models (Figures 11 and 12). The RandQReph strategy also shows improvements, with the best-performing strategy achieving 7.0% for GPT-3.5-turbo, 8.5% for GPT-4o-mini, 6.5% for Llama-3.1-8B-Instruct, and 13.7% for Qwen2.5-7B-Instruct relative improvements in Pass@10 on HumanEval (Figure 13). Back-translation also yields a 5.7% relative gain (Figure 14).

The paper also explores additional aspects: the effect of model sizes (Table 3) shows that diversified sampling provides larger gains for weaker models than for stronger ones; cost-performance tradeoffs (Figure 15) show that task-level perturbations use roughly the same number of output tokens as the baseline while solving more problems, and query-level perturbations with larger query batches (10 or 50 ideas per query) outperform the 1-idea setting; and a comparison with prompt optimization (Figure 16) shows that the diversified sampling Single variant attains higher problem-solving accuracy than Dipper while consuming fewer total output tokens.

In conclusion, the paper states: "this work provides a systematic analysis of the effect of sampling diversity. We offered a theoretical perspective on why exploration diversity enhances Best-of-N performance. Building on our main theorem, we derive a diversity–fidelity trade-off principle that guides the design of sampling strategies that introduce diversity while preserving fidelity. Following this guidance, we instantiate a set of perturbation styles. We theoretically and empirically analyze when diversified exploration remains effective, showing that sampling diversity remain effective across a wide range of conditions, while its benefits still depend on the thinker model's strength, the perturbation cardinality, the policy model's strength, the computational budget, and how many perturbations are generated per query. Our theoretical proposition shows that majority voting constitutes a failure mode in which diversity does not lead to performance gains, a behavior that is also confirmed empirically, suggesting that diversity cannot be applied indiscriminately."

Improvements for AI systems

Based on the paper, here are specific improvements to AI systems and their resulting capabilities:

Improvement: Implement a perturbation-aware sampling module that automatically generates and injects task-aligned, query-level perturbations (e.g., solution ideas, rephrasings) into prompts before each sampling attempt, replacing naive i.i.d. sampling from a single prompt.

Capabilities:

  • Achieve up to 10.8% relative improvement in EM@100 on MMLU-Pro, 8.2% on MATH, and 4.7% on HumanEval over direct sampling (GPT-4o-mini, budget of 100 solutions).

  • Break out of local solution clusters by exploring multiple diverse reasoning/coding paths, leading to faster convergence of Best-of-N as N grows.

  • Automatically select between Single (self-generated), Dual (separate thinker model), and Diverse (multi-model pool) perturbation strategies based on available compute.

Abstract

Large language model (LLM) scaling inference is key to unlocking greater performance, and leveraging diversity has proven an effective way to enhance it. Motivated by the observed relationship between solution accuracy and meaningful response diversity, we systematically study the effect of prompt diversity in scaling inference. We theoretically explain why diversified sampling improves Best-of- N scaling, showing that responses generated from diverse prompts after Best-of- N selection exhibit significantly lower error rates than those produced from stationary prompts. Building on this analysis, we derive a diversity-fidelity trade-off principle, that guides the design of sampling strategies introducing diversity. From this guidance, we instantiate a family of effective perturbation styles. We theoretically and empirically characterize when diversified exploration remains effective, demonstrating that it works under a variety of conditions, and we further show that under majority voting, diversity may vanish. Finally, we systematically evaluate the effectiveness of sampling diversity and show that, when applied appropriately in different contexts, meaningful perturbations yield stronger, task-dependent gains as diversity increases. Overall, this work provides a systematic analysis that offers a theoretical and empirical foundation for understanding how sampling diversity affects LLM inference-time scaling.

Sources

Related papers