The exponential distribution of the order of demonstrative, numeral, adjective and noun
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "The exponential distribution of the order of demonstrative, numeral, adjective and noun".
Tom: The frequency of preferred orders for noun phrases formed by demonstrative, numeral, adjective,
Jane: First, who's behind it and why it matters.
Paper summary: Tom: Now we’re moving into segment two where we get a clearer picture of what they actually did in "The exponential distribution of the order of demonstrative, numeral, adjective and noun." They set out to investigate whether an exponential or a power law distribution better models the frequency rank distribution for the twenty-four possible orders <ref:2502.06342#pg0,an exponential or a power law distribution>.
Jane: The core thesis here is that they found an exponential distribution provides a much better fit than the power law model when analyzing the actual data from Dryer datasets. This finding challenges previous assumptions about how these linguistic patterns are distributed in reality.
Lu: Lu sees this as a fundamental shift in how we interpret syntactic frequency data; it moves the focus away from assuming universal power-law behavior and toward identifying simpler, exponential probabilistic models.
Meng: Meng is wondering if they focused on just one specific subset of these orders or if they analyzed all twenty-four possibilities in detail to get this result.
Lalam: Lalam thinks the study's importance lies in its direct test against established hypotheses, showing that the exponential model actually performs statistically better than the power law approximation for this specific linguistic task.
Tom: The paper does lay out that they analyzed the frequency of rank r, where rank one is the most frequent order, and they defined two primary mathematical approximations: a power law represented by f(r) = cr-alpha, and an exponential distribution described by f(r) = ce-beta r.
Jane: That’s right Tom; they are directly comparing these two mathematical forms to see which one matches the observed frequency of rank r in the data. The paper sets up this comparison very clearly using Cysouw’s two thousand ten proposal against Martin et al <ref:2502.06342#pg0>.’s two thousand twenty assumptions <ref:2502.06342#pg0>.
Lu: Lu thinks the mathematical definition they use, like defining F x as F x = r max times r x, is a clever way to translate discrete linguistic frequencies into a continuous distribution framework for testing purposes.
Meng: From an engineering view, Meng wants to know if the method of translating these discrete counts into these specific formulas allows for any kind of scalable application beyond this specific noun phrase order analysis.
Lalam: Lalam believes that the methodology is strong because it’s rigorous; they didn't just guess which distribution was right; they used information-theoretic tools like AICc and BIC to make a formal selection based on predictive performance and model parsimony.
Tom: And those formal criteria are what give weight to their conclusion, showing that the geometric distributions were superior when we pitted Geometric one against Geometric two using those specific metrics <ref:2502.06342#pg1>.
Jane: So, in essence, they didn't just look at a graph and guess; they used rigorous statistical comparison to determine that the geometric distribution was the best fit for modeling this rank frequency data.
Lu: Lu sees this as demonstrating how computational linguistics can move from qualitative observation to quantitative proof by systematically testing competing probabilistic hypotheses against empirical data sets.
Meng: Meng wonders if the complexity of setting up these specific truncated distributions is something that would be too heavy to implement in a real-time language processing pipeline without significant optimization.
Lalam: Lalam thinks the rigor is what matters most here because it validates that when we look at linguistic patterns, we need precise statistical tools to move past mere intuition about which curve looks 'better'.
Conclusion: Tom: We’re wrapping up with the conclusion of "The exponential distribution of the order of demonstrative, numeral, adjective and noun," and it really boils down to a few key points for our listeners. They confirm that the geometric distribution is the best model for this specific linguistic phenomenon.
Jane: That means that while power laws are still relevant in some areas like word frequencies, they aren't necessarily the dominant statistical description for preferred word order patterns involving demonstratives, numerals, adjectives, and nouns in this context.
Lu: Lu thinks this supports a broader view that simpler models often capture the essence of complex linguistic phenomena better than overly complex ones forced onto the data.
Meng: Meng is thinking about how this might simplify our design philosophy; if we can rely on an exponential model for these specific structures, it could lead to less resource-intensive computational solutions.
Lalam: Lalam feels this finding suggests that there are no hard constraints limiting word order variation in languages, but rather that some orders just aren't attested because of undersampling, which is a more realistic picture.
Tom: So the takeaway here is that the paper strongly suggests reality is simpler than previously assumed, indicating no hard constraint on word order variation when we look at these specific phrase structures.
Jane: It’s a key point for understanding how language evolves; it implies that linguistic rules might be more about probability and sampling density than absolute prohibitions on certain sequences existing.
Lu: Lu sees this as an exciting direction for generative AI; if we can model the structure probabilistically without hard constraints, we open up possibilities where the AI explores a wider, more natural space of possibilities.
Meng: Meng thinks that this statistical insight could actually make our language models less brittle when encountering novel or slightly unexpected phrasing because they won't be locked into an overly strict set of rules.
Lalam: Lalam believes this paper challenges the inevitability of power laws in linguistics, suggesting that empirical generalizability is better served by simpler exponential structures.
Tom: That’s it for our discussion on "The exponential distribution of the order of demonstrative, numeral, adjective and noun," and we hope this gives our listeners a clearer picture of this interesting statistical work.
Ramon Ferrer-i-Cancho
Quantitative, Mathematical and Computational Linguistics Research Group · Department de Ci`encies de la Computaci´o · Universitat Polit`ecnica de Catalunya
cs.CL, physics.soc-ph
Submitted: 2025-02-10
Updated: 2026-01-22
Comments: minor corrections (typos and English errors)
DOI: 10.1080/09296174.2026.2617705
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 73/100
The gist: The frequency of preferred orders for noun phrases formed by demonstrative, numeral, adjective, and noun has been investigated to determine if an exponential or power law distribution better models
Key concepts
- Exponential Distribution
- This is a type of probability distribution where the likelihood of an event decreases steadily and continuously as the rank increases. The study found this model accurately describes how often specific word orders occur, outperforming power laws.
- Power Law Distribution
- A power law suggests that some events (like very frequent word orders) happen much more often than others, following a specific mathematical relationship. The research found this model is not the best fit for noun phrase orders in this linguistic context.
- Word Order Rank
- This refers to the frequency rank of a preferred word order, where rank 1 is the most frequent order and subsequent ranks are less frequent. Researchers analyzed how these ranks are distributed across all possible combinations.
- Geometric Distribution
- A discrete distribution that models the probability of observing a certain number of occurrences before some stopping condition is met. This model was found to be superior to power laws for describing the actual frequency of preferred noun phrase orders.
Terminology
Summary
The frequency of preferred orders for noun phrases formed by demonstrative, numeral, adjective, and noun has been investigated to determine if an exponential or power law distribution better models their actual distribution. This research finds that an exponential distribution is a much better model than a power law distribution for this linguistic phenomenon and suggests that there are no hard constraints on word order variation.
The Core Finding
We find that an exponential distribution is a much better model.
This finding challenges the view that power-law distributions, such as Zipf’s law for word frequencies, are inevitable. The study investigates the actual distribution of the 24 possible orders and finds that an exponential model provides a superior fit compared to power laws.
Methodology and Data
The researchers investigated the frequency of rank (r) of a preferred order, defined such that rank 1 is the most frequent order. They utilized data from two datasets: Dryer (2006) and Dryer (2018), which provide frequencies in languages, genera, and adjusted numbers of languages. The analysis involved defining statistical properties such as the total frequency (F0), the average frequency rank (⟨r⟩ = F1/F0), and the maximum frequency rank (rmax). The research then translated informal terminology into specific discrete distributions, including right-truncated zeta distributions (Zeta 1 and Zeta 2) and right-truncated geometric distributions (Geometric 1 and Geometric 2), as the variable (rank) is discrete.
Model Comparison
The study compared four specific models: Zeta 1, Zeta 2, Geometric 1, and Geometric 2. The models were mathematically defined based on their parameters—such as the exponent α for zeta distributions or q for geometric distributions—and theoretical constraints (e.g., R ≤ N=24). Model selection was performed using information-theoretic criteria: the corrected Akaike Information criterion (AICc) and the Bayesian Information Criterion (BIC). The best model was determined by minimizing a criterion, with AICc being favored when predicting future observations and BIC being more useful in selecting a correct model due to its stronger penalty for lack of parsimony.
Results and Interpretation
Visual diagnostics across normal, linear-log, and double logarithmic scales indicated that while a power law might appear straight on the linear-log scale, the exponential distribution is better supported by curves showing a straight line with a negative slope in the log-log scale. Model selection confirmed that geometric distributions were superior to power laws. Furthermore, when comparing Geometric 1 and Geometric 2 using information criteria:
AICc provides more support for Geometric 2 while BIC provides more support for Geometric 1.
The study concluded that the absence of a hard constraint on word order variation is more likely than its presence, suggesting that unattested orders merely result from undersampling, consistent with Cysouw’s view.
Implications for Linguistic Laws
The findings challenge the inevitability of power laws like Zipf’s law in linguistics. The abundance of exponential-like distributions has implications for the debate on linguistic laws, as evidence is found across various systems including SOV structures and primate vocalizations. The conclusion is that there is no empirical support for the claim that power laws are inevitable, even in a linguistic context. This suggests that researchers may fail to see simpler underlying structures when examining elementary exponential distributions with a few parameters.
Conclusion
The geometric distribution, as a discrete exponential-like distribution, is the best model for the distribution of preferred orders in the noun phrase. The analysis indicates that there is no statistically robust evidence for a hard constraint limiting word order variation in languages; rather, it suggests that all orders are a priori possible but some are not attested due to undersampling. The absence of a hard constraint is more likely than its presence when prioritizing empirical generalizability and theoretical consistency (the virtue of BIC). This finding challenges the inevitability of power laws in linguistics. The best model for the underlying distribution is Geometric 1, which aligns with the hypothesis that all orders are possible but some are not attested due to undersampling. The use of a score with a stronger penalty for lack of parsimony (BIC) leads to this conclusion, indicating that the absence of a hard constraint is more likely than its presence. This suggests that reality may be simple, but researchers may fail to see it.
The gist: An exponential distribution is a much better model than a power law distribution for the frequency rank distribution of noun phrase orders, and the evidence suggests there are no hard constraints on word order variation.
How it works
-
The research investigates the actual distribution of 24 possible orders using data from Dryer (2006) and Dryer (2018).
-
It compares two primary distributions: a power law, approximated by a zeta distribution, and an exponential-like distribution, approximated by a geometric distribution.
Improvements for AI systems
As a diligent researcher, I have analyzed the provided paper, The exponential distribution of the order of demonstrative, numeral, adjective and noun,
written by Ramon Ferrer-i-Cancho. The core scientific finding is that word order frequency distributions in this specific linguistic context are better modeled by an exponential/geometric distribution rather than a power law (like Zipf's law).
Based on this evidence, here are the specific improvements for AI systems:
-
Acknowledge and Prioritize Exponential/Geometric Modeling over Power Law Assumptions in Linguistic Sequence Prediction:
-
Implement Rank-Based Probability Modeling using Geometric Distributions for Syntactic Structure Generation:
-
Develop Robust Model Selection Techniques prioritizing Generalizability (BIC) over Predictive Fit (AICc) for Constraint Discovery:
-
Incorporate Rank-Truncation Constraints derived from the Geometric Model to Simulate
Hard Constraints
in Language Generation:
Specific Improvements and Capabilities:
-
The AI system can be improved by shifting its generative or predictive models away from purely power-law assumptions (which often lead to overestimation of low-frequency events) toward distributions characterized by a geometric decay function, which the paper identifies as superior for noun phrase order frequencies.
-
The AI system can specifically utilize the parameters derived from the best-fit Geometric 1 or Geometric 2 models (specifically the parameter 'q') to predict the probability of observing any specific word order rank in a sentence construction task, rather than relying on a universal power-law exponent.
-
By prioritizing BIC (Bayesian Information Criterion) scoring over AICc (Akaike Information Criterion), the AI system can be improved to be more robust against overfitting on finite datasets. This is crucial for discovering
hard constraints
—i.e., identifying structural rules that limit word order variation—because BIC is shown to favor the simpler model (Geometric 1, suggesting no hard constraint) when consistency and generalizability are prioritized. -
The AI system can be engineered to explicitly test the hypothesis of hard constraints by comparing the likelihoods of Geometric 1 (no constraint) versus Geometric 2 (hard constraint, where only ranks up to R are possible). The paper suggests that if Geometric 1 is preferred by BIC, it supports the theory that word order variation is not strictly constrained but rather a function of undersampling.
In summary, the improved AI system will be better at modeling linguistic variability by using an exponential decay framework and will be more reliable in distinguishing between actual structural constraints (hard constraints) and mere sampling artifacts (undersampling).
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering