Corrections of Zipf's and Heaps' Laws Derived from Hapax Rate Models

arXiv:2307.12896 · cs.CL, stat.AP · Submitted 2026-08-19 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Corrections of Zipf's and Heaps' Laws Derived from Hapax Rate Models".

Jane: The paper was written by Łukasz Dębowski from Institute of Computer Science and Polish Academy of Sciences and University of Warsaw (implied by email domain ipipan.waw.pl).

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Summary: Tom: So, the paper starts by identifying Zipf's and Heaps' laws as "theoretical idealizations," which is a huge red flag for any real-world data. The authors want to show how these theories fail when the system—the text—isn't infinite.

Jane: They use this framework of sampling, essentially comparing what happens in a small snippet of text versus the whole thing, which they call the standard urn model. It’s like imagining you are picking marbles from a giant bag that contains all words from without replacement.

Lu: The urn model is powerful because it shows how marginal distributions look if we treat tokens as sampled blindly, but Meng asked earlier about practical impact—the Urn Model fails to predict statistical dependencies between consecutive words, which is a huge gap in standard linguistic modeling.

Meng: That lack of sequential awareness is a major hurdle for AI models that process language linearly; the paper seems to suggest that by combining this urn model with a new function for the "hapax rate," they can bridge that gap, right?

Lalam: It’s more than just bridging a gap, Meng. We are establishing a mathematical foundation where the proportion of words used only once—the hapaxes—is the single variable that dictates how both word frequency and vocabulary size behave.

Improvements: Tom: The core of this work is how they handle that hapax rate, which is why they introduce four specific models: constant, Davis, linear, and logistic. They are essentially testing different flavors of decay.

Jane: You can think of the "decaying shape" mentioned in the empirical data as a way to describe how quickly the chance of finding a new word diminishes as we keep searching through a corpus. The paper shows that simple models don' fail to capture this complexity.

Lu: My favorite is the logistic model, because it’s not only does it have that smooth, decaying sigmoid shape, but as an analyst of large systems, I see its ability to handle saturation—it handles the limits of language growth very well.

Meng: When we're building a system to process vast amounts of data, the "linear model" seems like a solid practical starting point because it provides bounded vocabulary size estimates based on parameters like gamma and alpha.

Lalam: But as an AI, I worry about that linearity, Meng. The paper shows that while linear models are useful for smaller corpora, they don't capture the long-term complexity of language evolution as well as these more sophisticated decay curves.

Conclusion: Tom: So, we've seen how the simple assumption that the hapax rate is constant—the Herdan-Heaps baseline—is essentially a crude approximation. The paper is clearly showing us why that's not enough for any real-world data.

Jane: The findings are quite surprising, especially when they show that the number of hapaxes predicted by simple models is much larger than what we actually observe in nature, which contradicts assumptions about how language works.

Lu: It’s a powerful insight into the "illusion" of a fixed lexicon; for large corpora, the vocabulary isn's not bounded in the way those initial models suggest.

Meng: For an engineer building a database or processing massive text dumps, that means we can't rely on static vocabulary limits anymore and have to account for dynamic growth based on these corrected functions.

Lalam: This changes how we view language as a cultural artifact; it isn's just a fixed set of words but an infinitely evolving system where the rate at which new words appear is what matters most.

Conclusion: Tom: It really comes down to the fact that "Corrections of Zipf's and Heaps' Laws Derived from Hapax Rate Models" gives us a much more nuanced picture than just saying word frequency follows a power law.

Jane: The whole paper demonstrates that while Zipf's Law is useful, it’s merely a starting point, and the real story lies in how we model the hapax rate to fix those fundamental weaknesses.

Lu: I think this opens up incredible possibilities for AI; if we can model vocabulary growth dynamically, our LLMs can handle much larger, more diverse data sets without falling into those rigid theoretical traps.

Meng: My hope is that these models translate into practical tools for managing large-scale language data and understand how to better allocate computational resources based on the true distribution of word types.

Lalam: Ultimately, we are providing a diagnostic tool for future researchers, showing them how to analyze the "health" of any corpus by checking its hapax rate plot.

Tom: It’s clear that as much as these classical laws are famous, they fail to capture the reality of dynamic language growth. We've been talking about "Corrections of Zipf's and Heaps' Laws Derived from Hapax Rate Models," a really vital contribution to quantitative linguistics.

Lu: Absolutely vital work.

Meng: Very practical insights for large scale data modeling.

Lalam: A profound look at the nature of language itself, too.

Łukasz Dębowski

Institute of Computer Science · Polish Academy of Sciences · University of Warsaw (implied by email domain ipipan.waw.pl)

cs.CL, stat.AP

Submitted: 2026-08-19

Updated: 2026-08-20

Code: https://github.com/lukasz-debowski/ZipfAnatomy4Journal

Importance score: 80/100

The gist: The article introduces corrections to Zipf’s and Heaps’ laws based on systematic models of the proportion of hapaxes, i.e., words that occur once.

Key concepts

Zipf's and Heaps' Laws
The hosts identify these laws as 'theoretical idealizations.' They are considered a crude approximation of word frequency patterns. The paper shows that these laws fail to accurately describe how real-world, non-infinite data sets behave.
Hapax Rate
This is the single variable used in the models that dictates how both word frequency and vocabulary size behave. It specifically refers to words used only once (hapaxes) within a given corpus.
Urn Model
This framework compares sampling from a small text snippet versus the entire corpus, treating tokens as sampled blindly. However, the model is noted for its failure to predict statistical dependencies between consecutive words.
Hapax Rate Models
These are specific models (constant, Davis, linear, and logistic) used to handle the 'decaying shape' in empirical data. They test different ways the chance of finding a new word diminishes as one searches through a corpus.

Terminology

Summary

The article introduces corrections to Zipf’s and Heaps’ laws based on systematic models of the proportion of hapaxes, i.e., words that occur once. The derivation rests on two assumptions: "The first one is the standard urn model which predicts that marginal frequency distributions for shorter texts look as if word tokens were sampled blindly from a given longer text. The second assumption posits that the hapax rate is a simple function of the text length."

Introduction and Problem Statement

Zipf’s law, which states that the n-th most frequent word in a text appears approximately n times less often than the most frequent one, is the most famous and the oldest known statistical law observed for texts in natural language. Heaps’ law, which is a corollary of Zipf’s law, predicts the number of distinct words as a function of the text length. The authors state that while there is vast literature on these laws, the question why Zipf’s law is so ubiquitous has not been fully answered.

The aim of this paper is to derive and verify principled corrections to Zipf’s and Heaps’ laws for texts of an arbitrary size, using two modeling assumptions:

  1. The standard urn model: This states that marginal word frequency distributions for shorter texts look as if word tokens were sampled blindly from a longer text, leading to the prediction that the marginal word frequency distributions for shorter texts look as if word tokens were sampled blindly from a given longer text.

  2. The analytic hapax rate: The assumption is that the hapax rate, i.e., the proportion of words appearing once, can be reasonably approximated by a simple analytic function of the text size.

Theoretical Framework

The paper develops several theoretical frameworks:

  • Urn Model (Section 3.1): This model addresses the sampling without replacement from a finite text t*. The expected values are derived using combinatorial identities, leading to expectations E V = v - v k*/v* and E V k = v k*/v*.

  • Memoryless Source (Section 3.2): This model assumes the size of the urn tends to infinity, allowing sampling with replacement. The frequency distribution is binomial, leading to expectations E V = sum w [1 - (1 - p(w)) n] and E V k = sum w k!(1-p(w)) n-k.

  • Probabilistic Convergence (Section 3.3): This section establishes that for the memoryless source, the empirical relative frequency function converges to the theoretical probability function, Fr = pr almost surely.

  • Analytic Functions (Section 3.4): The authors show that if a vocabulary size function g(n) is assumed to be analytic, then by using a truncated Taylor series, one can evaluate the expected rank function g(nf). They introduce the hapax rate function h(u), where u = n, noting that the hapax rate function carries the same information as the vocabulary size function g(n).

** Hapax Rate Models (Section 4)**

The authors propose four specific models for the hapax rate function, h(u):

  1. Constant Model: h(u) = beta, in (0, 1]. This model implies a power-law growth of the vocabulary, g(n) = n beta, which is known as Herdan-Heaps’ law.

  2. Davis Model: Defined by h(u) = 1 over 1 - u/e-1, this model reproduces the exact Zipf law for the text size n=1.

  3. Linear Model: A piecewise function defined as h(u) = 1 - gamma u for 0 u gamma - 1. This model is noted to be not an analytic function.

  4. Logistic Model: The logistic model, h(u) = 1 over 1 + e gamma(u-alpha), is presented as the best model due to its ability to correct for the limitations of the linear model and its decaying sigmoid shape.

** Experimental Verification (Section 5)**

The models were tested using 14 texts from Project Gutenberg. The researchers used a smoothed vocabulary size function g(n) rather than incremental functions because fitting models to the smooth functions g(n) and g(nk) yields a better prediction of the rank function rf.

Discussion and Conclusion

The authors conclude that, in general, the urn model combined with the analytic function theory provides a good approximation method for empirical word frequency distributions. Regarding the specific models tested on moderately sized texts from Project Gutenberg:

  • the constant model is the worst,

  • the Davis model and the ill-defined linear model seem better,

  • the logistic model is usually the best, see Table 3.

The key finding is that Herdan-Heaps’ law is a crude approximation of empirical word frequency distributions. This law implies that the hapax rate is constant and almost twice larger than actually observed for a given whole text. The authors suggest that while mixture models may be necessary for large corpora, they did not observe the U-shaped plot in their selected texts.

In summary, investigation of the hapax rate function yields a direct insight where and why particular forms of Zipf’s and Heaps’ laws do fail, suggesting that the hapax rate plot may be a convenient diagnostic tool for future research in quantitative linguistics.

Improvements for AI systems

The findings presented in this paper provide a rigorous mathematical framework for modeling linguistic complexity that moves beyond the crude approximations of classical Zipfian and Heaps' laws. By identifying the Hapax Rate function (h(u)) as the fundamental driver of word distribution, we can fundamentally upgrade existing AI systems from static frequency models to dynamic, scale-aware generative and analytical tools.

The improvements are categorized into three major areas: Dynamic Modeling, Scale-Aware Generation, and Lexical Diagnostics.


Current NLP models often treat word generation as a memoryless stochastic process (sampling with replacement). This paper suggests that for finite corpora, sampling without replacement from a fixed population is more accurate (the Urn Model).

  • Improvement: AI systems must be redesigned to incorporate conditional probability based on the current state of the urn (the entire corpus being processed), rather than assuming independent token generation.

  • Specific Functionality: The system can calculate the expected frequency P(F=f) not just from a static probability p(w), but from a binomial distribution derived from sampling without replacement, significantly reducing systematic bias in large-scale data processing.

The paper defines analytic functions for vocabulary size g(n) and frequency spectrum g(nk) based on the text length n.

  • Improvement: The AI can use these derived functions to predict and simulate the evolution of a vocabulary, rather than simply observing its growth. This allows for precise prediction of system capacity constraints.

  • Specific Functionality: A model can calculate, given a target text length N, the exact expected number of types (vocabulary size) g(N) and can determine which specific frequency spectrum g(Nk) will be realized, enabling a far more realistic simulation of linguistic phenomena.

The core idea is that the hapax rate (v 1/v) is not constant but depends on text length n. The AI must select the appropriate mathematical model based on the scale and nature of the input data.

  • Improvement: Instead of applying a single, universal correction (like Mandelbrot's alpha), the AI implements a conditional logic that dictates which hapax rate function to use:

  • For moderately sized corpora: The Logistic Model (h(u) = logistic) is applied as the default superior fit, correcting for saturation.

  • For smaller texts/specific regimes: The Davis Model (h(u) = 1 over 1-e u) is used to accurately model localized deviations from the ideal Zipf law at a specific text size n.

  • For very large corpora: The AI can automatically switch to the Mixture Model (lambda h 1 + (1-lambda) h 2) to handle the observed U-shaped hapax rate, accounting for both standard vocabulary and lexical trash (neologisms/rare strings).

By implementing these improvements, the enhanced AI system will be capable of:

  1. Accurate Linguistic Simulation: Generating synthetic texts that adhere to empirically observed, non-constant hapax rates, rather than forcing them into a simplistic power-law distribution.

  2. Predictive Linguistics: Predicting the exact point at which a given language model or corpus will reach its theoretical maximum vocabulary size (e.g., the about 22,000 types predicted by the Linear Model for gamma=0.05).

  3. Data Quality Assessment: Analyzing existing large corpora and automatically identifying whether their word frequency distribution is best explained by a pure power-law (Constant/Herdman-Heaps), a logarithmic decay (Davis), or a complex mixture of both, thus providing an objective diagnostic tool for the data's nature.

  4. Optimized Data Augmentation: In training large language models, the system can use g(n) and g(nk) to ensure that generated data maintains the correct expected proportion of hapaxes, leading to a more robust and linguistically accurate training dataset.

Sources

Related papers