Ask a Language Model for Lottery Numbers: Concentration in Repeated Six-of-49 Outputs

arXiv:2610.00052 · cs.IR, cs.AI, cs.CL · Submitted 2026-09-04 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Ask a Language Model for Lottery Numbers".

Jane: Six language-model configurations were evaluated for their ability to generate six distinct random integers from 1 to 49,

Tom: First, who's behind it and why it matters.

Title and authors: Tom: So, to get into Segment Two, we're going to talk about the title and authors of this paper, "Ask a Language Model for Lottery Numbers: Concentration in Repeated Six-of-forty-nine Outputs," which sets the stage for what we're about to discuss.

Jane: That paper clearly lays out the scope right away by focusing on the specific task—asking an AI for six distinct random integers from one to forty-nine—and who conducted this evaluation.

Lu: The authors are Dmitrij Żatuchin from the Department of Information Technologies at Estonian Entrepreneurship University of Applied Sciences, and they're working alongside Rankfor.AI in Tallinn, Estonia on this work.

Meng: I'm just curious about the specific focus of their evaluation; are they testing a single model or a whole family?

Lalam: They evaluate six different language-model configurations in total, which gives us a broad view of how this concentration manifests across different architectures.

Tom: That breadth is important because it tells us that this isn't just an artifact of one particular model; it’s a general behavior we need to be aware of across the board.

Jane: It’s interesting because the title itself immediately signals that the main concern isn't just "can it generate numbers," but rather "is what it generates truly random?"

Lu: And they make a strong point by evaluating this against simulated uniform sampling thresholds, which is a solid way to ground their findings in established statistical expectations.

Meng: Grounding it in simulation helps show the reader exactly how far off these results are from the expected behavior of independent uniform six-of-forty-nine sampling.

Lalam: So, what we're seeing is that the models, even when fresh contexts are used for every call, still lean toward repeated outputs instead of truly independent draws.

Tom: Exactly; it’s a fundamental observation about how these language models are operating under their default configurations on this specific kind of request.

Jane: And the authors make it clear that they aren't trying to pinpoint the exact cause right away, which is smart because it prevents them from getting bogged down in technical details prematurely.

Lu: That choice to keep the mechanism open for now allows them to focus on quantifying the effect first, which is a very methodical approach.

Meng: It means we're getting data on *what* happens, even if we don't have the complete picture of *why* it happens yet.

Lalam: And this sets us up perfectly for the next part where we look at what those quantified effects actually look like in practice.

The paper's summary: Tom: Now, let’s look at the actual summary section of "Ask a Language Model for Lottery Numbers: Concentration in Repeated Six-of-forty-nine Outputs" to understand the main findings they reported.

Jane: The summary explains that they tested six different language models on requests for six distinct random integers from one to forty-nine, and it highlights that the results show substantial concentration in repeated ticket outputs that significantly deviates from independent uniform sampling.

Lu: Specifically, the paper reports a key finding: across all tested deployment settings, the modal unordered ticket accounted for between twenty-two point five percent and sixty-eight point zero percent of all valid responses.

Meng: That range is quite wide, Tom; it means the concentration isn't uniform across every single model they tested; some models are much more repetitive than others.

Lalam: It’s not just a slight dip or rise; it's a substantial tendency toward output concentration rather than genuine randomness when we look at the overall distribution of answers.

Tom: That is what they emphasize, and it’s a pretty big finding because it moves the conversation away from whether these models *can* do the task to how well they actually do it in terms of quality.

Jane: It means that relying on these models for tasks where true randomness is critical might be risky if we don't account for this tendency toward repetition.

Lu: The comparison they made with the physical lottery samples, which showed effective diversities of forty-one point one and forty-one point four at smaller sample sizes, provides a real benchmark against which to measure these findings.

Meng: So when you look at the raw data, we see that most of the engines are performing worse than those archived samples when it comes to achieving true diversity in their number selections.

Lalam: It really shows that even with large sample sizes, we aren't seeing the ideal distribution curve they expect from independent uniform six-of-forty-nine sampling.

Tom: That leads us into the next part where we discuss what these results actually mean for us in terms of practical implications for how we use these AI systems.

The paper's improvements: Jane: Moving into the section on improvements, the authors are suggesting that future work should focus on creating better methods to handle this concentration in their generation process.

Lu: They propose developing mechanisms to detect and mitigate severe "concentration" of outputs toward a small subset of possibilities when generating diverse samples from a defined range like one to forty-nine.

Meng: I think that’s where we need to focus our engineering efforts; building in an internal check that monitors the Effective Diversity, Neff, and flags outputs if they fall below a certain threshold would be really useful.

Lalam: That proactive detection strategy feels like a necessary step because instead of just accepting the output as it is, we should have something built to push the generation process back toward more uniform distribution.

Tom: They also suggest developing prompt engineering strategies that can dynamically adjust the model's behavior based on what the task requires or what concentration metrics are observed during generation.

Jane: That suggests a sophisticated layer of control where we aren't just picking one static prompt, but maybe switching between variants depending on the desired output quality.

Lu: If we can map out which prompt phrasing yields better diversity across different models, that would give us a lot more leverage in designing better applications for these systems.

Meng: Integrating diversity scoring as a primary metric during model selection would be a powerful way to steer development toward models that are inherently more robust in terms of output variety.

Lalam: Prioritizing those metrics over just raw speed or accuracy is the right direction because for tasks requiring randomness, variety matters far more than how fast the response comes back.

Tom: And they also touch on differentiating between true randomness and these biased stochastic processes, which is a crucial conceptual step for anyone trying to truly understand what these AI are capable of doing.

Jane: That kind of distinction would help us provide users with a much clearer understanding of the inherent randomness, or lack thereof, in the AI's output.

Conclusion: Tom: So we’ve covered a lot about the paper, and now it’s time for our wrap-up as we discuss the conclusion of "Ask a Language Model for Lottery Numbers: Concentration in Repeated Six-of-forty-nine Outputs."

Jane: To summarize, the main implication is that these models show substantial concentration in repeated ticket outputs, which is incompatible with independent uniform selection of six distinct numbers from one to forty-nine.

Lu: This suggests that we need better ways to handle this tendency toward repetition if we want reliable results in any application involving random number generation.

Meng: From an engineering view, it means any system relying on this kind of generation needs a diversity check built in before those outputs are used for critical decisions.

Lalam: It’s about acknowledging that the output isn't always perfectly random, which is a fair and necessary caution when using these tools.

Tom: It’s a cautionary tale about what happens when we don't properly account for these internal statistical biases in language model generation, as shown in "Ask a Language Model for Lottery Numbers: Concentration in Repeated Six-of-forty-nine Outputs."

Jane: We’ve learned that while the paper doesn't pinpoint the exact cause, we do have a clear indication of where the concentration is occurring and how it varies across different models and prompts.

Lu: The future work they suggest about prompt testing is a valuable path forward for developing more sophisticated tools that can actively steer these systems toward better diversity.

Meng: I think integrating diversity scoring into model selection pipelines is a practical step toward building more reliable AI components for generation tasks.

Lalam: Ultimately, the paper reminds us that we need to be careful with our expectations when asking an AI to perform tasks that demand true randomness, because the output isn't guaranteed to be uniform.

Dmitrij Żatuchin

Department of Information Technologies, Estonian Entrepreneurship University of Applied Sciences (EUAS), Tallinn, Estonia · Rankfor.AI

cs.IR, cs.AI, cs.CL

Submitted: 2026-09-04

Updated: 2026-09-04

Comments: 6 pages, 1 figure. Data, code, and collector at github.com/Rankfor/rankfor-open (research/lotto-models)

Code: https://github.com/Rankfor/rankfor-open

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 83/100

The gist: Six language-model configurations were evaluated for their ability to generate six distinct random integers from 1 to 49, revealing substantial concentration in repeated ticket outputs that deviates

Key concepts

Neff
This metric measures how diverse a model's output is regarding the numbers generated. A higher Neff score indicates better diversity, while a lower score suggests more repetition in the numbers returned. It helps quantify the concentration of specific number sets.
Modal Unordered Ticket
A ticket is defined by any six unique numbers regardless of order. The modal unordered ticket is the single most frequently occurring set of six numbers across all valid responses from a model. Analyzing this reveals which specific combination the model tends to favor.
Uniform Benchmark
This represents the ideal scenario where every possible combination of six distinct numbers from 1 to 49 has an equal chance of being selected. The study compares the models' actual output concentration against this perfect, unbiased distribution to determine if they are truly random.

Terminology

Summary

Six language-model configurations were evaluated for their ability to generate six distinct random integers from 1 to 49, revealing substantial concentration in repeated ticket outputs that deviates significantly from independent uniform sampling. This study demonstrates that across tested deployment settings, the modal unordered ticket accounted for between 22.5% and 68.0% of valid responses, indicating a tendency toward output concentration rather than true randomness.

Design and Benchmarks

The experiment involved evaluating six engines—gpt-5.6-luna, claude-sonnet-5, gemini-3.7-flash, grok-4.5, mistral-large, and Perplexity sonar—on requests for six distinct random integers from 1–49 across four English prompt variants. A total of 200 attempted calls were distributed evenly across these engines and prompts (50 calls per variant). Every call opened a fresh context, and the provider-default configurations used by the collector were evaluated without systematic variation in decoding settings. The results are scored by the effective diversity of number frequencies, defined as Neff = 1/P49 i=1 s 2 i, where si is the share of all returned numbers landing on i. A source using all 49 numbers equally scores 49, while a source always returning the same six scores 6.

Evaluation Metrics and Comparison

The primary metric for concentration is Neff. Across the six engines tested, Neff ranged from 9.9 to 18.0 against simulated fifth-percentile floors of 46.6 to 46.7 at their respective sample sizes, with a pMC value of 0.00005 each when no simulated dataset was as concentrated as the observed results. The two archived Polish Lotto samples provided a physical-lottery comparison, yielding effective diversities of 41.1 and 41.4 at smaller sample sizes, which are compatible with the uniform benchmark at their sample sizes. The test is Monte Carlo against the uniform ticket sampler, where pMC = (1 + neff ≤ obs) / (B + 1), one-sided with ties counted.

Whole Ticket Behavior

Analysis of whole answers demonstrated that stronger results were found at the level of complete unordered tickets. Tickets are canonicalised as sorted six-number sets, meaning two answers with the same six numbers in different orders count as one ticket. For example, gpt-5.6-luna produced 8 distinct tickets in 200 valid answers and returned its most frequent one (7, 14, 22, 31, 38, 46) in a modal share of over 50.5%. Similarly, sonar returned the same six numbers in a modal share of 68.0% of its answers. The smallest number found in every system’s modal unordered ticket was consistently 7.

Prompt and Ordering Effects

The performance varied significantly based on the prompt used. For instance, gpt-5.6-luna returned only 2 to 3 distinct tickets with a 96% modal share under the en base and en lottery phrasings. Conversely, grok-4.5 showed a dramatic shift from an 8% modal share (Neff 23.1) under the base phrasing to a 72% modal share (Neff 7.4) when the word lottery appeared in the prompt (en lottery). Furthermore, most systems returned numbers in ascending order, with gemini-3.7-flash and gpt-5.6-luna returning the six numbers already sorted in 100% of parsed answers.

Scope and Limitations

The study rejects independent uniform six-of-49 sampling under the tested configurations, asserting that a biased stochastic process is still random. The experiment does not distinguish between memorisation, decoding behaviour, prompt conditioning, retrieval, request routing, or response caching as the cause of concentration. Specifically regarding tool configurations: sonar is retrieval-enabled by default while the other five received no search parameters; therefore, no single shared mechanism should be inferred across them. The physical baseline comes from operator statistics for two same-season samples 27 years apart (1999 and 2026), and the analysis covers engine outputs alone, not repeated whole physical tickets.

Conclusion

Across the six tested configurations, the modal unordered ticket accounted for between 22.5% and 68.0% of valid responses, which is incompatible with independent uniform selection of six distinct numbers from 1-49. Run-to-run variation is not evidence of coverage; every engine put between 22.5% and 68.0% of its answers on one ticket.

Improvements for AI systems

Based on the provided scientific paper, here are specific improvements that could be implemented in AI systems, along with what those improved systems could achieve:


  1. Improve output reliability for constrained random generation tasks (e.g., lottery number selection, unique ID generation).

  2. Implement a mechanism to detect and mitigate severe concentration of outputs toward a small subset of possibilities when generating diverse samples from a defined range (like 1-49).

  3. Develop prompt engineering strategies that dynamically adjust the model's behavior (e.g., increasing output diversity) based on the task requirements or observed output concentration metrics.

  4. Integrate diversity scoring as a primary metric during model selection, specifically prioritizing models that demonstrate effective diversity of number frequencies over raw accuracy or speed.

  5. Create a system that can differentiate between true randomness (uniform sampling) and biased stochastic processes inherent in LLM generation, allowing for better quantification of output quality under different deployment settings (temperature, prompt variants).

  6. Establish standardized auditing protocols for repeated queries to systematically map how model behavior changes across different prompting strategies and tool configurations, ensuring that observed concentration is not due to caching or retrieval mechanisms.

These improvements would enable the following capabilities:

  1. An AI system could reliably generate a set of six distinct, truly diverse random numbers from a constrained range (e.g., 1-49) with high statistical confidence, minimizing the risk of generating highly predictable or non-random results that might be favored by biased sampling techniques.

  2. A system could proactively identify and flag model outputs exhibiting excessive concentration (low Effective Diversity, Neff) toward a few specific outcomes, allowing downstream applications to reject or re-sample those results before they are used in high-stakes decision-making processes (e.g., financial modeling or critical sampling).

  3. The system could intelligently select the optimal prompt variant for a given generation task; for instance, it could automatically switch from a base prompt to a lottery variant if the goal is explicitly lottery number selection, leveraging empirical data on which phrasing yields higher diversity.

  4. A model selection pipeline would prioritize models based on their measured ability to maintain broad coverage of the output space, rather than just overall performance metrics, leading to more robust and less prone-to-failure AI components.

  5. The system could provide a confidence score indicating whether the generated result is likely derived from uniform sampling or a concentrated pattern, helping users understand the inherent randomness (or lack thereof) in LLM outputs.

  6. An auditing framework would allow researchers to systematically test if output concentration is caused by internal model behavior (memorization/decoding) versus external factors like caching or retrieval mechanisms, leading to more trustworthy and reproducible AI systems.

Sources

Related papers