The Shrinking Landscape of Linguistic Diversity in the Age of Large Language Models

arXiv:2502.11266 · cs.CL · Submitted 2025-02-16 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "The Shrinking Landscape of Linguistic Diversity in the Age of Large Language Models".

Jane: The paper was written by Zhivar Sourati, Farzan Karimi-Malekabadi, Meltem Ozcan, Colin McDaniel, Alireza Ziabari et al. from Department of Computer Science, University of Southern California. and Center for Computational Language Sciences, University of Southern California. and Department of Psychology, University of Southern California. and Information Science Institute, University of Southern California..

Tom: Stay tuned as we take you through the paper and discuss its implications.

Paper discussion segment 1: Tom: We’ve seen how the authors identified this issue, and now they’re diving into the actual findings of "The Shrinking Landscape of Linguistic Diversity in the Age of Large Language Models," which is really showing a trend of homogenization.

Jane: Homogenization here means that even if people are writing different stories or articles, they are all starting to sound more and more alike, which is a massive loss for unique voices.

Tom: It’s not just about the content being the same as when AI rewrites it; the way you phrase things is what's changing to make it consistent.

Lu: The finding that this variance is decreasing across different writing styles suggests that LLMs are pushing us toward some kind of linguistic average, which feels very constrained for my theoretical framework.

Meng: And I’m seeing this in the data; the reduction in complexity variance is a measurable shift, and it’s consistent whether we look at creative prompts or academic papers.

Lalam: It's a major concern because if everyone starts using the same predictable language, it could limit how we express our specific cultural experiences.

Tom: The paper says this happens even though the core meaning of the text remains true to be true, which is important because that means AI isn't changing *what* is said, but it’s certainly changing *how* it’s said.

Jane: We're seeing a pattern where the AI-written texts tend to look more standardized than human-written ones across the board.

Lu: I think this trend is especially noticeable in places like academic writing where LLMs are becoming a powerful tool for consistency, but that consistency is itself becoming a constraint.

Meng: That’s what I'm worried about; if it’s standardizing academic writing, it might be suppressing novel ways of thinking about problems.

Lalam: And we need to make sure that when we use AI tools for clarity, we aren't inadvertently losing the unique texture of human expression.

Paper discussion segment 2: Tom: Now, the authors are moving into more specific experiments with "The Shrinking Landscape of Linguistic Diversity in the Age of Large Language Models," looking at how these models affect our ability to read about people’s personal traits.

Jane: This is where it gets really interesting because linguistic patterns are supposed to be a signature—they tell us things about the person, even if they don're just writing an essay or a social media post.

Tom: The study shows that when these specific stylistic cues are smoothed out by AI, our ability to infer traits like personality or political affiliation drops significantly.

Lu: This is huge for me because it directly challenges how we map linguistic patterns onto psychological constructs in the first place.

Meng: And I think this has practical implications for systems that try to assess candidates based on their written work, as those markers are becoming less reliable.

Lalam: If we can't rely on the subtle ways people structure their thoughts, it means we risk losing a lot of the nuance that defines our individual identities.

Tom: The researchers found that even though LLMs retain the meaning, they systematically erase these unique linguistic signals needed to judge traits like empathy or openness.

Jane: It’s like trying to read a highly detailed portrait and having all the fine brushstrokes smoothed away by sanding it down too much.

Lu: And this erosion of predictive power is something we need to study further in terms how we measure what's actually meaningful in language.

Meng: From an engineering view, if our classifiers are suddenly getting less reliable because of the input style, that means a the algorithms that rely on those markers could fail or be biased.

Lalam: We have to find a way to ensure that as AI assists us with clarity and structure, we are not sacrificing the depth of character and personality in our communication.

Paper discussion segment 3: Tom: The authors are also looking at how these effects change based on the specific tools or prompts you use, which is really showing a lot of nuance in "The Shrinking Landscape of Linguistic Diversity in the Age of Large Language Models."

Jane: It’s not just that *any* LLM causes this; we're seeing differences between Gemini and Llama three for example, so it isn't a one-size-fits-all problem.

Tom: They found that certain tools are more "preservative" of those unique traits than others, which is a huge variable we need to track.

Lu: This variation suggests that the alignment process itself—how the model is trained—is influencing how it preserves or alters specific linguistic structures.

Meng: And I think this detail matters for us because if you’re building an AI assistant, knowing which LLMs tend to be more "conservative" of stylistic elements allows us to build better checks into our workflows.

Lalam: It means that the way we choose our writing assistants can have a direct impact on how much personality shines through in our work.

Tom: The paper suggests that while the trend is toward homogeneity, some connections—like those tied to morality or empathy—are surprisingly resilient and hold up better than others.

Jane: That’s an important distinction; it seems certain aspects of human expression are more robust than others when the AI gets involved.

Lu: I wonder if those resilient traits are the ones that are less reliant on subtle syntactic variation and more tied to specific lexical choices?

Meng: A lot of that sounds like it has to do with which specific features the LLM is trained to value over general stylistic flow.

Lalam: We need to be mindful of this—that if we prioritize certain types of content, we might unintentionally boost those traits while losing others that are more subtle or complex.

Conclusion: Tom: So, after looking at the data and the experiments in "The Shrinking Landscape of Linguistic Diversity in the Age of Large Language Models," it’s clear that this is a serious issue we need to address.

Jane: It's not just that AI is making our writing look cleaner; it's subtly erasing the signals that tell us who we are and how we think.

Tom: We’ve seen this across Reddit, academic papers, and news articles, showing it’s a widespread phenomenon.

Lu: I think this suggests a long-term trend where the cognitive diversity of our collective output is at risk if we aren't careful.

Meng: From an engineering perspective, we need to build ways to detect and compensate for this drift in predictive power as AI becomes more prevalent.

Lalam: We must ensure that as we leverage these powerful tools, the rich tapestry of human identity remains intact and isn't replaced by a uniform style.

Tom: Before we go, I want to get a final thought from each of you on "The Shrinking Landscape of Linguistic Diversity in the Age of Large Language Models."

Lu: I hope that we can find ways to use these models without sacrificing the unique voices they are meant to assist.

Meng: We have a responsibility to make sure that our tools aren' are designed not just for efficiency, but for reliability and preserving human input.

Lalam: And it's about finding a better balance between empowering AI and protecting the world’s diverse way of speaking.

Tom: Thank you all so much; it’s been a truly thought-provoking discussion on "The Shrinking Landscape of Linguistic Diversity in the Age of Large Language Models."

Jane: I agree; we'll be thinking about this for a long time to come. Lu, Meng, and Lalam: (Short interjections of agreement).

Zhivar Sourati, Farzan Karimi-Malekabadi, Meltem Ozcan, Colin McDaniel, Alireza Ziabari, Jackson Trager, Ala Tak, Meng Chen, Fred Morstatter,, Morteza Dehghani

Department of Computer Science, University of Southern California. · Center for Computational Language Sciences, University of Southern California. · Department of Psychology, University of Southern California. · Information Science Institute, University of Southern California.

cs.CL

Submitted: 2025-02-16

Updated: 2026-08-24

Code: https://github.com/ArthurHeitmann/arctic

Importance score: 72/100

The gist: This paper investigates how the widespread adoption of large language models (LLMs) as writing assistants is linked to "notable declines in linguistic diversity" and may interfere with the societal

Key concepts

Linguistic Homogenization
This is the trend where diverse writings begin to sound increasingly alike, regardless of content. The hosts discuss how this standardization represents a massive loss for unique voices and stylistic variance in human communication.
Inferring Traits from Language
This refers to the ability to deduce personal characteristics, such as personality or political affiliation, based on subtle linguistic patterns in someone's writing. The discussion notes that LLMs can systematically erase these unique signals, making assessment less reliable.
Variance Reduction
This finding suggests that LLMs are pushing language toward a measurable 'linguistic average.' It is described as a reduction in complexity variance, indicating that the AI constrains the natural variability and uniqueness of human expression.

Terminology

Summary

This paper investigates how the widespread adoption of large language models (LLMs) as writing assistants is linked to notable declines in linguistic diversity and may interfere with the societal and psychological insights provided by language. As LLMs are designed to generate the most statistically likely continuation of a text, they risk promoting a flattening of linguistic expression that prioritizes conformity over individuality.

Observational trends in writing complexity

Study 1 examines temporal trends in writing styles using historical data from three diverse sources:

  1. Reddit (r/WritingPrompts) creative stories.

  2. Patch News localized community articles.

  3. arXiv Preprints in Computer Science and Linguistics/Vision categories.

By measuring the variance of features such as the Vocabulary Simpson Index, Shannon Entropy, and Type-Token Ratio, the researchers identified a consistent decline in the variance of writing complexity following the introduction of ChatGPT. This suggests that LLMs are influencing communication by promoting a convergence toward stylistic norms and reducing variability in both creative and scientific expression.

Experimental evidence of homogenization

Study 2 experimentally replicates these findings by prompting GPT-3.5, Llama 3 70B, and Gemini Pro to rewrite original texts using Syntax Grammar or Rephrase prompts. The results demonstrate that while the core content of texts is retained, LLMs significantly reduce the variability in writing complexity. Although semantic similarity remained high—with 87% of scores above 0.95—the models' revisions caused a significant reduction in linguistic diversity. This confirms that LLMs act as tools that emphasize conformity over individuality while preserving the primary information intended by the author.

Erosion of predictive power for personal traits

Study 3 explores whether this homogenization obscures crucial linguistic markers essential for identifying individual characteristics. Researchers trained classifiers on original texts to predict demographic and psychological attributes, then applied them to LLM-rewritten versions. They found an average 6% decline in the absolute F1 score across various constructs, including:

  • Demographic attributes (age, gender, and political affiliation).

  • Personality dimensions (the Big Five model).

  • Dispositional empathy and moral foundations.

The study revealed that LLM-rewritten texts exhibit greater linguistic homogeneity, which skews predictions toward specific profiles. For example, LLM-generated text was more likely to be associated with authors being predicted as male, older, or politically liberal.

Disruption of established lexical associations

Study 4 utilizes a top-down approach to examine how LLMs affect the relationship between fine-grained lexical cues and personal traits. Using validated lexicons like LIWC and the NRC Emotion Lexicon, the researchers found that many well-established associations between these cues and personal traits were washed away by LLM involvement. For instance, the link between gender and negative emotion words was largely erased, and associations between extraversion and social or pronoun usage were significantly weakened. This indicates that LLMs do not merely introduce noise but systematically alter the connection between language and identity, potentially undermining efforts in clinical psychology, mental healthcare, and personnel selection.

Improvements for AI systems

1. Stylistic Entropy-Preserving Decoding Algorithms

  • Improvement: Modify the decoding process (e.g., top-p, temperature, or nucleus sampling) by integrating a real-time Complexity Variance Penalty. This involves calculating the input text's linguistic complexity features—specifically Type-Token Ratio (TTR), Vocabulary Simpson Index, and Hapax Legomena—and penalizing any output that significantly reduces the variance of these metrics.

  • Capability: The AI can polish grammar, syntax, and spelling while mathematically ensuring it does not smooth out the original author's unique lexical richness or syntactic complexity.

2. Identity-Signature Rewriting (ISR) Modules

  • Improvement: Implement a pre-processing Signature Detection layer using lightweight classifiers to identify the demographic and psychological linguistic markers (e.g., pronoun frequency, emotional valence, and social referents) of the input text. This signature is then used as a set of soft constraints or via Adapter layers during the generative process to maintain those specific statistical distributions.

  • Capability: An AI writing assistant that can rephrase or summarize text while preserving the author's linguistic fingerprint, ensuring that gendered speech patterns, personality-driven word choices (e.g., extraversion/neuroticism cues), and cultural dialects remain detectable and intact.

3. Counter-Homogenization Alignment (CHA) in RLHF

  • Improvement: Reformulate the Reinforcement Learning from Human Feedback (RLHF) reward function to include a Diversity Reward. Instead of only optimizing for helpfulness and harmlessness—which current research shows biases models toward a polite, agreeable, and politically liberal persona—the model is rewarded for maintaining the stylistic distance between different user inputs.

  • Capability: The AI avoids the flattening effect where all users' writing eventually converges into a single, idealized persona, thereby preventing the systematic erasure of marginalized linguistic styles and diverse cultural markers.

4. Diagnostic-Sensitive Generative Guardrails

  • Improvement: Develop context-aware Feature-Guardrails that trigger when an LLM detects text in sensitive domains (e.g., mental health, clinical psychology, or legal testimony). These guardrails prevent the model from suppressing non-fluent markers—such as specific hesitation markers, affective word frequencies, or cognitive processing cues—that are statistically vital for diagnostic accuracy.

  • Capability: An AI used in clinical or professional settings that can assist in drafting communications without inadvertently masking the subtle linguistic indicators of psychological distress, cognitive impairment, or specific emotional states required for early medical diagnosis.

5. Demographic-Invariant Style Controllers

  • Improvement: Integrate Style-Control parameters into the model architecture that allow users to explicitly set a Preservation Coefficient. This coefficient would act as a weight between the model's statistical likelihood (the most probable next word) and the input text's historical distribution of demographic-linked features (e.g., age-related future-focus words or gendered social markers).

  • Capability: A highly granular control system that allows professional editors, researchers, or authors to decide exactly how much of their original voice is sacrificed for the sake of correctness, preventing the unintended shift toward a dominant demographic profile (e.g., older, male-coded language).

Sources

Related papers