Rank-Turbulence Delta and Interpretable Approaches to Stylometric Delta Metrics

arXiv:2604.19499 · cs.CL · Submitted 2026-04-21 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "Rank-Turbulence Delta and Interpretable Approaches to Stylometric Delta Metrics".

Tom: This article introduces two novel measures for authorship attribution—Rank-Turbulence Delta and Jensen–Shannon Delta—which generalize Burrows’s classical Delta by employing distance functions derived from probabilistic distributions,

Jane: First, who's behind it and why it matters.

Paper summary: Tom: Essentially, this paper is proposing Rank-Turbulence Delta and Jensen–Shannon Delta as new tools for authorship attribution because they generalize Burrows’s classical Delta by using distance functions based on probabilistic distributions. The authors claim these methods offer a more interpretable framework for stylistic analysis, which is a big deal when you need to understand *why* two texts are different.

Jane: They start by re-casting uncentred word-frequency vectors as probability distributions, which then lets them apply distance measures from information theory and complex systems analysis instead of just using the standard Manhattan distance on z-scored vectors, which is what Burrows’s Delta typically uses.

Lu: The paper goes into developing probabilistic distance measures, like Jensen–Shannon divergence and Rank-Turbulence divergence, and they also introduce a token-level decomposition that makes every Delta distance numerically interpretable by showing the "token-level contributions."

Meng: That token-level decomposition sounds promising for practical implementation because it lets researchers pinpoint exactly which individual words are driving the stylistic variation between texts.

Lalam: I think that ability to see the specific tokens contributing allows for a much deeper level of textual understanding, moving beyond just saying "Text A is different from Text B."

Tom: Exactly. They show that Rank-Turbulence Delta achieves attribution accuracy comparable to Cosine Delta, and Jensen–Shannon Delta consistently matches or even exceeds the performance of canonical Burrows’s Delta. The authors test this across four literary corpora in English, German, French and Russian.

Jane: Their experimental validation is quite thorough; they assess clustering quality and attribution performance on these diverse datasets while also testing the robustness of their methods under temporal and stylistic variation using the SOCIOLIT corpus.

Lu: The results suggest that Rank-Turbulence Delta shows consistently high performance across languages and word frequency ranges, which points toward a stable representation of stylistic structure in these metrics.

Conclusion: Tom: So, wrapping up this discussion on "Rank-Turbulence Delta and Interpretable Approaches to Stylometric Delta Metrics," the main point is that these new metrics offer a structured quantitative map of lexical contrasts, which helps guide qualitative analysis after the initial statistical comparison.

Jane: The authors have really expanded on Burrows’s original concept by introducing probabilistic measures that allow for a clearer look at how word frequencies actually contribute to stylistic separation in a way that's more transparent.

Lu: What I find particularly interesting is the finding about mid-frequency words; the study suggests these play a particularly important role in stylistic discrimination, especially when you look at lower values of the rank parameter alpha.

Meng: From an engineering standpoint, knowing which specific tokens are driving that separation means we can build systems that are more sensitive to those key lexical signals without just treating all word frequencies equally.

Lalam: If this framework helps us understand the asymmetry between lexical presence and absence, it could help us design AI models that better capture nuanced cultural expression rather than just surface-level vocabulary counts.

Tom: It really shows that neither exclusively dominant nor exclusively rare items are sufficient to capture stylistic identity; mid-frequency vocabulary carries substantial signal in these new approaches. This has big implications for how we model human creativity and language use.

Jane: Ultimately, the implication is a more sophisticated way to analyze authorship attribution because it provides not just an accuracy score, but also a pathway to understanding the underlying linguistic structure driving that score.

Dmitry Pronin, Evgeny Kazartsev

HSE University

cs.CL

Submitted: 2026-04-21

Updated: 2026-10-02

Comments: Published in Digital Scholarship in the Humanities. The version of record is available at https://academic.oup.com/dsh/advance-article-abstract/doi/10.1093/llc/fqag072/8692587 Code available at: https://github.com/DDPronin/Rank-Turbulence-Delta

Journal ref: Digital Scholarship in the Humanities, 2026

DOI: 10.1093/llc/fqag072

Code: https://github.com/DDPronin/Rank-Turbulence-Delta

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 77/100

The gist: This article introduces two novel measures for authorship attribution—Rank-Turbulence Delta and Jensen–Shannon Delta—which generalize Burrows’s classical Delta by employing distance functions

Key concepts

Rank-Turbulence Delta
This is a novel authorship attribution metric based on comparing word frequency profiles using a divergence measure derived from rank-based representations. It uses a formula that incorporates token ranks and a tuning parameter to compare the hierarchical ordering of words between texts, allowing for more nuanced stylistic comparisons.
Jensen–Shannon Delta
This is another new metric that compares two texts based on their probability distributions. It measures the divergence between two distributions by comparing each text to their weighted average, where weights sum to one. It is built upon the Kullback–Leibler divergence and provides a robust way to compare stylistic profiles.
Token-Level Decomposition
This technique breaks down the final Delta distance calculation to show exactly which individual words contribute most significantly to the difference between two texts. This helps researchers visualize which specific vocabulary items are driving the stylistic distinction, making the results more meaningful.
Probabilistic Distance Measures
These are mathematical tools used to quantify how different two word frequency distributions are. Instead of simple subtraction, they use information theory concepts like divergence to measure stylistic distance based on how likely certain words are to appear in each text.

Terminology

Summary

This article introduces two novel measures for authorship attribution—Rank-Turbulence Delta and Jensen–Shannon Delta—which generalize Burrows’s classical Delta by employing distance functions derived from probabilistic distributions, thereby providing a more interpretable framework for stylistic analysis.

The gist: Rank-Turbulence Delta attains attribution accuracy comparable with Cosine Delta; Jensen–Shannon Delta consistently matches or exceeds the performance of canonical Burrows’s Delta.

Theoretical Foundation and Vector Representation

The study begins by re-casting uncentred word-frequency vectors as probability distributions, which allows for the application of distance measures from information theory and complex-systems analysis. The canonical Burrows’s Delta compares standardized (z-scored) word-frequency profiles across texts, defined by the Manhattan (L1) distance between these z-vectors. However, uncentred z-vectors can be normalized to represent probability distributions where each component is the relative weight of a token within the stylistic profile of the text. This transformation opens the possibility of employing divergence measures such as Jensen–Shannon divergence and Rank-Turbulence divergence.

Probabilistic Distance Measures

The paper develops probabilistic distance measures based on these probability distributions. The Jensen–Shannon divergence compares two texts not directly to each other, but to their weighted average, where the mixture weights are chosen such that the sum equals one. The Kullback–Leibler divergence is also introduced as an element of the Jensen–Shannon divergence. Furthermore, for rank-based representations, the rank-turbulence divergence is proposed to compare hierarchical ordering of tokens using a formula involving ranks and a tuning parameter, where smaller values of alpha amplify the contribution of lower-ranked tokens.

Token-Level Decomposition and Interpretation

A systematic token-level decomposition is developed to render every Delta distance numerically interpretable by identifying the token-level contributions. For canonical Burrows’s Delta, the contribution of a token is defined as:

(12) (1) (2) p p    = −

This decomposition allows researchers to visualize which individual words exert the greatest influence on stylistic differences between texts. For Cosine Delta, contributions arise from products rather than sums and can be positive or negative, with specific interpretations based on whether components exceed or fall below their corpus means. Similarly, for Jensen–Shannon Delta and Rank-Turbulence Delta, token contributions are extracted in a manner analogous to the canonical Burrows’s Delta formula.

Experimental Validation and Robustness

The effectiveness of these methods is assessed across four literary corpora in English, German, French, and Russian. Experiments on clustering quality demonstrate that Rank-Turbulence and Cosine Delta show consistently high performance across languages and mfw ranges, suggesting stable representations of stylistic structure. The study also tests the robustness of these measures under temporal and stylistic variation using the SOCIOLIT corpus. Robustness checks include:

  1. Variation of most frequent words (mfw) size, showing that Jaccard overlap between adjacent mfw settings remained consistently substantial across metrics.

  2. Document-level bootstrap resampling, where the mean Jaccard overlap of the top-30 tokens ranged between approximately 0.37 and 0.50, indicating a stable signal rather than sampling noise.

  3. A removal experiment confirming that removing top-contributing tokens led to a consistent reduction in inter-author distance across all four metrics, confirming their functional role in document separation.

Discussion on Stylistic Significance

The findings suggest that mid-frequency words play a particularly important role in stylistic discrimination, as clustering quality is maximized at lower values of the rank parameter alpha. The research further reveals an asymmetry between lexical presence and absence: while centred representations show strongly negative deviations tend to occupy lower positions in the rank hierarchy, the reorganization of prominent lexical signals appears more structurally informative for document separation. This indicates that mid-frequency vocabulary carries substantial signal, reinforcing that neither exclusively dominant nor exclusively rare items suffice to capture stylistic identity. The proposed framework provides a structured quantitative map of lexical contrasts that guides subsequent qualitative analysis.

Limitations and Future Directions

Limitations include the context-dependent nature of parameter selection, specifically for alpha, which may require corpus-specific stability analysis. While the methods cover multiple languages and corpora, further testing on heterogeneous historical datasets is suggested. Additionally, the transformation required for Rank-Turbulence divergence involves an explicit reordering of features with asymptotic complexity comparable to sorting operations. The proposed interpretative framework is noted as a tool to guide qualitative analysis rather than replacing close reading.

Data and Code Availability

Code and data used in this study are publicly available at https://github.com/DDPronin/Rank-Turbulence-Delta, with metadata and derived representations provided in the repository. The preprocessing pipeline involved standardizing term-frequency matrices using the 20,000 most frequent word types across corpora, followed by computing both centred and uncentred z-standardized representations.

Improvements for AI systems

As a fastidious and diligent researcher, I have analyzed this paper, RANK-TURBULENCE DELTA AND INTERPRETABLE APPROACHES TO STYLOMETRIC DELTA METRICS. The core contribution of this work lies in moving beyond purely predictive authorship attribution towards stylometric analysis that prioritizes transparency and interpretability.

Here are the specific improvements to AI systems that can be made using the methodologies described in this paper, along with what those improved systems can achieve:


  1. Improvement: Integration of Token-Level Stylistic Fingerprinting via Rank-Turbulence Delta (RTD)

  2. Improvement: Implementation of a Token Contribution Mapping layer for any text classification or attribution model trained on stylometric features.

  3. What the Improved AI System Can Do:

4.1. Identify and quantify the specific lexical items (words or n-grams) that are driving a particular textual similarity score (e.g., between two texts, an author and their corpus). This moves the system from merely stating Text A is similar to Text B to explaining Text A is similar to Text B because of the high contribution of tokens X, Y, and Z.

4.2. Enable targeted stylistic probing: Researchers can use this mapping to test hypotheses about authorial style (e.g., Does the difference between Author 1 and Author 2 hinge on the usage frequency of 'pacing' verbs vs. adverbs?).

4.3. Enhance interpretability in complex attribution tasks: For unsupervised or exploratory tasks where high predictive accuracy isn't the sole goal, this layer provides a transparent bridge between high-dimensional vector distances and meaningful linguistic features, facilitating close reading and validation against philological intuition (as noted in Section 1).

  1. Improvement: Adoption of Probabilistic Distance Metrics (Jensen–Shannon Delta) for Similarity Assessment

  2. Improvement: Replace standard Euclidean or Cosine distance measures with Jensen–Shannon Divergence (JSD) when comparing text distributions, especially in contexts involving varying corpus sizes or when the underlying feature space is not perfectly Gaussian.

  3. What the Improved AI System Can Do:

4.1. Provide a more robust measure of stylistic similarity that is intrinsically based on probability theory (Section 3). JSD measures how much information would be lost if one text were approximated by a mixture distribution of the two texts, offering a symmetric and information-theoretic view of divergence (Equation 8).

4.2. Maintain performance stability across different data regimes: By using probabilistic representations derived from uncentred z-vectors, the system maintains structural integrity even when dealing with non-negative distributions, which is crucial for consistency across multilingual corpora (English, German, French, Russian).

  1. Improvement: Dynamic Feature Weighting via Rank-Turbulence Parameter Tuning

  2. Improvement: Incorporate a tunable parameter (like the rank sensitivity parameter α) into the distance calculation framework to allow the model to dynamically shift its focus between high-frequency dominant features and subtle mid-frequency signals.

  3. What the Improved AI System Can Do:

5.1. Optimize performance for specific stylistic nuances: The system can be tuned to prioritize structural vocabulary (high frequency, low α) or distinctive lexical signals (mid-frequency, moderate α). This allows the AI to adapt its sensitivity based on the research question—whether it seeks broad stylistic tendencies or idiosyncratic markers.

5.2. Achieve stable clustering and classification: Experimental results show that optimal performance is often found at intermediate values of α, suggesting a more nuanced understanding of authorial contrast than relying solely on extreme high-frequency items, leading to more robust clustering solutions (Section 10).

  1. Improvement: Robustness Checks for Feature Selection (mfw)

  2. Improvement: Implement automated robustness checks that assess the stability of the top-contributing tokens against minor perturbations in vocabulary size (mfw) and document sampling variation (bootstrap resampling).

  3. What the Improved AI System Can Do:

6.1. Ensure reliable feature selection: The system will not rely on a single, narrowly defined set of most frequent words. By confirming that the top-contributing tokens remain stable across varying mfw sizes (as shown in Appendix C), the resulting attribution or clustering decisions are proven to be based on a stable structural property of the authorial contrast, not an artifact of arbitrary vocabulary cutoff choices.

6.2. Increase confidence in attribution results: The system can report a confidence score based on the Jaccard overlap with adjacent mfw settings, providing assurance that the identified fingerprint tokens are genuinely discriminative and not sampling noise (Section 19).

In summary, these improvements transform a stylometric AI from a simple distance calculator into an advanced, explainable analytical tool capable of identifying, quantifying, and interpreting the precise lexical components responsible for stylistic variance.

Related papers