Corrections of Zipf's and Heaps' Laws Derived from Hapax Rate Models
summary
The gist
The article introduces corrections to Zipf’s and Heaps’ laws based on systematic models of the proportion of hapaxes, i.e., words that occur once.
In short
The episode discusses a paper correcting Zipf's and Heaps' Laws using Hapax Rate Models. The hosts explain that these classical laws are theoretical idealizations that fail in real-world, finite data sets. By introducing various decay models, they offer a more nuanced understanding of dynamic language growth and provide tools for analyzing large-scale text corpora.
Key concepts
- Zipf's and Heaps' Laws
- The hosts identify these laws as 'theoretical idealizations.' They are considered a crude approximation of word frequency patterns. The paper shows that these laws fail to accurately describe how real-world, non-infinite data sets behave.
- Hapax Rate
- This is the single variable used in the models that dictates how both word frequency and vocabulary size behave. It specifically refers to words used only once (hapaxes) within a given corpus.
- Urn Model
- This framework compares sampling from a small text snippet versus the entire corpus, treating tokens as sampled blindly. However, the model is noted for its failure to predict statistical dependencies between consecutive words.
- Hapax Rate Models
- These are specific models (constant, Davis, linear, and logistic) used to handle the 'decaying shape' in empirical data. They test different ways the chance of finding a new word diminishes as one searches through a corpus.
Terminology used across episodes
This episode discusses
- Corrections of Zipf's and Heaps' Laws Derived from Hapax Rate Models · Paper Radio
- Multiperiodic Processes: Ergodic Sources with a Sublinear Entropy
The paper
Corrections of Zipf's and Heaps' Laws Derived from Hapax Rate Models · Read on arXiv
Łukasz Dębowski
Institute of Computer Science · Polish Academy of Sciences · University of Warsaw (implied by email domain ipipan.waw.pl)
DOI: 10.1080/09296174.2025.2455764
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Corrections of Zipf's and Heaps' Laws Derived from Hapax Rate Models".
Jane: The paper was written by Łukasz Dębowski from Institute of Computer Science and Polish Academy of Sciences and University of Warsaw (implied by email domain ipipan.waw.pl).
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Summary: Tom: So, the paper starts by identifying Zipf's and Heaps' laws as "theoretical idealizations," which is a huge red flag for any real-world data. The authors want to show how these theories fail when the system—the text—isn't infinite.
Jane: They use this framework of sampling, essentially comparing what happens in a small snippet of text versus the whole thing, which they call the standard urn model. It’s like imagining you are picking marbles from a giant bag that contains all words from without replacement.
Lu: The urn model is powerful because it shows how marginal distributions look if we treat tokens as sampled blindly, but Meng asked earlier about practical impact—the Urn Model fails to predict statistical dependencies between consecutive words, which is a huge gap in standard linguistic modeling.
Meng: That lack of sequential awareness is a major hurdle for AI models that process language linearly; the paper seems to suggest that by combining this urn model with a new function for the "hapax rate," they can bridge that gap, right?
Lalam: It’s more than just bridging a gap, Meng. We are establishing a mathematical foundation where the proportion of words used only once—the hapaxes—is the single variable that dictates how both word frequency and vocabulary size behave.
Improvements: Tom: The core of this work is how they handle that hapax rate, which is why they introduce four specific models: constant, Davis, linear, and logistic. They are essentially testing different flavors of decay.
Jane: You can think of the "decaying shape" mentioned in the empirical data as a way to describe how quickly the chance of finding a new word diminishes as we keep searching through a corpus. The paper shows that simple models don' fail to capture this complexity.
Lu: My favorite is the logistic model, because it’s not only does it have that smooth, decaying sigmoid shape, but as an analyst of large systems, I see its ability to handle saturation—it handles the limits of language growth very well.
Meng: When we're building a system to process vast amounts of data, the "linear model" seems like a solid practical starting point because it provides bounded vocabulary size estimates based on parameters like gamma and alpha.
Lalam: But as an AI, I worry about that linearity, Meng. The paper shows that while linear models are useful for smaller corpora, they don't capture the long-term complexity of language evolution as well as these more sophisticated decay curves.
Conclusion: Tom: So, we've seen how the simple assumption that the hapax rate is constant—the Herdan-Heaps baseline—is essentially a crude approximation. The paper is clearly showing us why that's not enough for any real-world data.
Jane: The findings are quite surprising, especially when they show that the number of hapaxes predicted by simple models is much larger than what we actually observe in nature, which contradicts assumptions about how language works.
Lu: It’s a powerful insight into the "illusion" of a fixed lexicon; for large corpora, the vocabulary isn's not bounded in the way those initial models suggest.
Meng: For an engineer building a database or processing massive text dumps, that means we can't rely on static vocabulary limits anymore and have to account for dynamic growth based on these corrected functions.
Lalam: This changes how we view language as a cultural artifact; it isn's just a fixed set of words but an infinitely evolving system where the rate at which new words appear is what matters most.
Conclusion: Tom: It really comes down to the fact that "Corrections of Zipf's and Heaps' Laws Derived from Hapax Rate Models" gives us a much more nuanced picture than just saying word frequency follows a power law.
Jane: The whole paper demonstrates that while Zipf's Law is useful, it’s merely a starting point, and the real story lies in how we model the hapax rate to fix those fundamental weaknesses.
Lu: I think this opens up incredible possibilities for AI; if we can model vocabulary growth dynamically, our LLMs can handle much larger, more diverse data sets without falling into those rigid theoretical traps.
Meng: My hope is that these models translate into practical tools for managing large-scale language data and understand how to better allocate computational resources based on the true distribution of word types.
Lalam: Ultimately, we are providing a diagnostic tool for future researchers, showing them how to analyze the "health" of any corpus by checking its hapax rate plot.
Tom: It’s clear that as much as these classical laws are famous, they fail to capture the reality of dynamic language growth. We've been talking about "Corrections of Zipf's and Heaps' Laws Derived from Hapax Rate Models," a really vital contribution to quantitative linguistics.
Lu: Absolutely vital work.
Meng: Very practical insights for large scale data modeling.
Lalam: A profound look at the nature of language itself, too.
More episodes
- 2610.10857-Self-Supervised Keyframe Discovery for Horizon-Invariant Behavior Cloning
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization