How Does "English (US)" Become the Default? Triangulating Structural Bias Towards American English Across the LLM Pipeline

arXiv:2604.04204 · cs.CL, cs.AI, cs.CY, cs.ET, cs.LG · Submitted 2026-04-05 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "How Does "English (US)" Become the Default? Triangulating Structural Bias Towards American English Across the LLM Pipeline".

Jane: Large language models (LLMs) increasingly deploy only limited language settings, most notably “English (US),” despite global diversity and colonial history,

Tom: First, who's behind it and why it matters.

Title and authors: Tom: So, we're talking about "How Does 'English (US)' Become the Default? Triangulating Structural Bias Towards American English Across the LLM Pipeline," and it sounds like this paper is systematically checking every step of how an AI model learns its language. It’s not just one thing causing the bias; it’s a whole chain of decisions from data collection to final generation.

Jane: Exactly, Tom; what's striking is that they use a method called DIALIGN to estimate dialectal alignment using distributional evidence, which is designed to capture various contrasts like grammar and style simultaneously. It’s trying to get a holistic view of the bias rather than just looking at one area in isolation.

Lu: The authors are constructing this curated corpus of one thousand eight hundred thirteen AmE–BrE variants specifically for this investigation, which is a massive undertaking to capture that diversity before applying their triangulation method across the three stages they outline.

Meng: From an engineering standpoint, having them formalize the research questions around pretraining corpora audits and tokenizer representations gives us concrete checkpoints to test our own model architectures against these specific points of failure.

Lalam: I see this as a necessary postcolonial framing because it suggests that the way English is standardized and curated by digital dominance creates an epistemic injustice where one dialect, American English, gets implicitly prioritized.

The paper's summary: Tom: The core of the paper is that they are triangulating evidence across three main stages—pretraining corpora audits, tokenizer representations, and generative preferences—to show how structural bias toward American English is introduced and amplified throughout the LLM development pipeline.

Jane: They found that in the pretraining data stage, there's a systematic skew toward AmE, especially with orthographic variants like color versus colour where AmE spellings dominate by margins above seventy percent. That immediately tells us the foundation is already tilted.

Lu: And then they move to the tokenizer stage, revealing that British English forms have higher fertility than their American counterparts, which points to less efficient tokenization for BrE and a gap of about eighteen point seven two percent in vocabulary-based differences.

Meng: That fertility metric is very tangible; it means the tokenizer isn't handling certain British words as efficiently as others, which translates directly into computational overhead or potentially less nuanced representation for those forms.

Lalam: It’s a really important finding because it shows that the problem isn't just in the model itself, but in how we tokenize the language before it even enters the model's core structure.

The paper's improvements: Tom: The authors suggest some really practical steps for fixing this, focusing on three areas: building dialect-sensitive corpus construction, injecting BrE tokens into base tokenizers using DIALIGN, and being extremely careful about adopting pretrained tokenizers without regional adaptation.

Jane: It seems like the authors are pushing for a proactive approach rather than just analyzing the problem; they want us to actively design systems that respect regional linguistic differences from the start.

Lu: The suggestion about using DIALIGN to inject BrE tokens into base tokenizers is ambitious, but it proposes a way to dynamically expand the vocabulary based on distributional evidence, which is a creative use of that training-free method.

Meng: For us engineers, the recommendation against blindly adopting pretrained tokenizers without adaptation is crucial; it tells us we can't just plug and play a tokenizer and assume fairness across dialects.

Lalam: I think the most profound improvement they suggest relates to how we approach the system itself—they are advocating for component-wise design recommendations to prevent linguistic homogenization in the broader AI deployment landscape.

Conclusion: Tom: So, to wrap up, this paper on "How Does 'English (US)' Become the Default? Triangulating Structural Bias Towards American English Across the LLM Pipeline" shows that dialectal skew isn't accidental; it’s structurally embedded in the pipeline through data choice, tokenization inefficiencies for British English forms, and generative defaults.

Jane: It really underscores the risk of linguistic homogenization if we don't actively work on building more inclusive systems; we need to be aware of how these models might enforce specific sociolinguistic norms in high-stakes environments.

Lu: I think the long-term vision here is that by understanding this pipeline structure, we can move toward truly dialect-sensitive corpus construction and tokenization design that reflects global English diversity rather than just one dominant standard.

Meng: Practically speaking, the implication for us is a need to build in those audit layers we discussed earlier so we can flag when our current models are showing those high AmE preferences before they go into widespread use.

Lalam: Ultimately, this paper gives us a roadmap to prevent epistemic injustice by ensuring that the underlying AI infrastructure respects the full range of English varieties, making the cultural impact of these systems much fairer for everyone involved.

cs.CL, cs.AI, cs.CY, cs.ET, cs.LG

Submitted: 2026-04-05

Updated: 2026-09-28

Comments: Preprint

Project page: https://www.statista.com/statistics/262946/most-common-languages-on-the-internet

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 80/100

The gist: Large language models (LLMs) increasingly deploy only limited language settings, most notably “English (US),” despite global diversity and colonial history, raising foundational questions about

Key concepts

Pretraining Corpora Audits
This involved systematically checking major datasets used to train LLMs for imbalances. Researchers quantified differences between American and British English by looking at word choices and alignment signals within these massive training sets. The results showed a consistent, statistically significant skew favoring American English.
Tokenizers and Fertility
Tokenizers break down text into smaller units (tokens). 'Fertility' measures how many subword tokens are needed to represent a word. The study found that British English forms require more tokens than American English, suggesting tokenization is less efficient for BrE, which impacts how the model processes different dialects.
DIALIGN Method
DIALIGN is a training-free method used to measure dialectal alignment. It analyzes n-grams (sequences of words) and calculates 'Signed Divergence' to predict which language variant a model will generate. This allowed researchers to quantify the generative preference for AmE versus BrE in model outputs.
Generative Default
This refers to the language that an LLM produces most often when given no specific instructions, or under a neutral condition. The study found that American English is the dominant generative default for these models, producing AmE outputs at rates between 65% and 80% with high confidence.

Terminology

Summary

Large language models (LLMs) increasingly deploy only limited language settings, most notably “English (US),” despite global diversity and colonial history, raising foundational questions about which variety of English LLMs implicitly prefer. This study investigates how geopolitical histories of data curation, digital dominance, and linguistic standardization shape the LLM development pipeline by triangulating evidence across pretraining corpora, tokenizers, and generative behaviors to reveal structural bias toward American English (AmE).

Research Objectives

The investigation is formalized through three core research questions designed to trace dialectal asymmetries across the LLM development pipeline:

  1. To what extent do large-scale pretraining corpora skew toward American over British English? This is addressed by conducting corpus-level audits of major LLM pretraining datasets to quantify dialectal imbalance using both lexical variants and broader alignment signals.

  2. How do regional tokenizers encode variants, and what does this reveal about dialectal representation? This involves examining subword-level disparities across tokenizers developed in American, European, Chinese, and postcolonial contexts by analyzing fertility, which is defined as the average number of subword tokens per word.

  3. Do LLMs exhibit generative preferences for AmE over BrE? This is evaluated by assessing dialectal preferences in model outputs under contextual prompts and estimating alignment across lexical, grammatical, structural, stylistic, and multi-word contrasts using the DIALIGN method.

Methodology: Triangulation Across the Pipeline

The analysis employs a three-stage triangulation to surface structural bias: (i) audits of pretraining corpora, (ii) tokenizer analyses of regional tokenizers, and (iii) generative evaluations of model outputs. To support this, two primary resources were constructed: a curated corpus of 1,813 parallel AmE–BrE lexical variants and the DIALIGN, a dynamic, training-free method for estimating dialectal alignment using distributional evidence. DIALIGN operates through four stages: n-gram Extraction (n=2 to 5), Frequency Lookup via the Google Books Ngram corpus, computation of Signed Divergence per n-gram (LR(g)), and final Aggregation and Normalization to yield alignment probabilities (PAmE, PBrE).

Findings from Pretraining Corpora Audits

Audits of six major open-access pretraining corpora reveal a systematic skew toward AmE. Using variant-specific token distributions, the study found that All datasets show a statistically significant skew toward AmE, with the skew being strongest for orthographic variants (e.g., color vs. colour), where AmE spellings dominate with margins above 70%. DIALIGN, which captures broader contrasts, likewise indicates a clear AmE preference, often with high confidence. This suggests that dialectal skew is structurally embedded in the pretraining data underlying modern LLMs.

Findings from Tokenizer Representations

The study examined how tokenizers encode variants using metrics like fertility and granularity. Results show that BrE forms exhibit higher fertility than their AmE counterparts, indicating less efficient tokenization. This disparity is larger for vocabulary-based differences, reaching a gap of ∆v = 18.72%. Furthermore, tokenizers developed outside the USA, such as those in Europe (Mistral) and China (DeepSeek), show better BrE coverage, while Gemma achieves the lowest overall fertility across both dialects. Granularity analysis further supports this: BrE variants are overrepresented in the 3+ token bin, indicating greater fragmentation.

Findings from Generative Preferences

In generative evaluations, DIALIGN demonstrated that AmE is the dominant generative default. Under a default English condition, most models produce "65–80% AmE outputs often with high confidence (> 0.80). Even when explicitly prompted with British English (en-GB), AmE persists, rarely dropping below 40%. The uptake of BrE is noted to be stronger in informal domains but limited in formal ones," reflecting structural biases shaped by pretraining data and tokenizer design.

Broader Implications and Recommendations

The findings suggest that dialectal skew reflects how pretraining data can embed broader cultural tendencies, potentially leading to linguistic homogenization and epistemic injustice. The paper motivates practical steps toward more inclusive technologies, including:

  1. Dialect-sensitive corpus construction, such as leveraging metadata from web-scale datasets to enrich coverage for World Englishes.

  2. Dialect-sensitive vocabulary extension via DIALIGN to inject BrE tokens into base tokenizers.

  3. Caution against blindly adopting pretrained tokenizers without regional adaptation, emphasizing that tokenization introduces unfairness between languages.

The study concludes that AmE is the entrenched generative default across LLMs, persisting even under BrE prompts, highlighting risks for users expecting BrE norms in institutional contexts.

Improvements for AI systems

Based on the systematic analysis presented in this paper, here are specific, actionable improvements for AI systems and what those improved systems can achieve:


)1. Improved Dialect-Aware Pretraining Data Curation (Addressing RQ1 & RQ2)

The core improvement is shifting from massive corpus ingestion to structurally balanced corpus engineering.

The AI system should incorporate a mandatory, automated audit layer that measures the dialectal skew of ingested data across orthographic and vocabulary dimensions before training commences.

This audit must utilize the DIALIGN metric (or an equivalent distributional evidence estimator) on sampled n-grams from major pretraining corpora (like C4 or Dolma).

Specific Actions:

  1. Implement a Dialectal Skew Score for every training corpus, quantifying the AmE/BrE imbalance based on the frequency distributions of lexical variants and structural n-grams (n=2 to 5).
  1. Prioritize data sources or apply synthetic augmentation only when the resulting synthetic data passes a DIALIGN parity check against target dialectal distributions.
  1. Implement Dialect-Aware Tokenization Priors: Before finalizing the tokenizer, use fertility metrics (from Table 3) to identify where AmE and BrE variants are underrepresented or over-fragmented in the subword vocabulary, allowing for targeted vocabulary expansion to inject missing BrE tokens.

Capability of Improved System:

The resulting LLM will possess a significantly reduced AmE default bias. It will exhibit more balanced performance across English dialects, leading to outputs that are less likely to enforce a single sociolinguistic norm. This ensures that the model’s internal representations of English are not structurally biased toward American conventions, preventing the propagation of epistemic injustice related to linguistic homogenization.

)2. Enhanced Dialect-Sensitive Tokenization Strategy (Addressing RQ2)

The system must move beyond generic tokenizers to context-aware segmentation.

The AI system should utilize a modular tokenizer that allows for dynamic vocabulary allocation based on the specific dialect being queried or generating, rather than relying on a single, monolithic tokenizer.

Specific Actions:

  1. Develop Dialect-Specific Token Maps: Create separate subword vocabularies or weighting schemes for AmE and BrE variants, informed by the fertility analysis (Table 3). For example, if a model is generating in an informal context (like ELI5), the tokenizer should prioritize subwords that efficiently represent BrE vocabulary (e.g., lift vs. elevator).
  1. Implement Granularity Control: Allow users or downstream systems to select the desired granularity (1-token, 2-token, or 3+-token) for tokenization during inference, dynamically adjusting the model's internal segmentation strategy to mitigate BrE over-segmentation.

Capability of Improved System:

The improved system will demonstrate superior efficiency and accuracy when processing dialectal input. It will handle BrE forms more efficiently (lower fertility/better granularity) and maintain better lexical fidelity, reducing latency and improving the model's ability to correctly segment complex, idiomatic BrE phrasing without resorting to excessive subword fragmentation.

)3. Generative Preference Alignment Layer (Addressing RQ3)

The model must be explicitly trained or fine-tuned to recognize and respect dialectal conditioning.

Integrate a Dialectal Alignment Module into the generation pipeline that uses DIALIGN as a real-time feedback mechanism to steer outputs toward the desired dialect, even when prompted with the alternative.

Specific Actions:

  1. Implement Real-Time DIALIGN Scoring: During inference, for every generated output segment (n=2 to 5), calculate the (PAmE, PBrE) alignment score using a lightweight version of DIALIGN and apply the lexicon boost factor (β).
  1. Contextual Correction Mechanism: If the input prompt specifies British English (en-GB), but the generated text yields a high PAmE score (> 70%), an automated post-processing step should be triggered to rephrase or replace AmE-leaning n-grams with their BrE counterparts, ensuring adherence to the explicit instruction.

Capability of Improved System:

The system will achieve superior Dialectal Fidelity. It will reliably produce outputs that strictly adhere to the requested dialect, even under conflicting contextual cues. This solves the problem where models default to AmE when prompted with BrE; instead, it provides a measurable, quantifiable mechanism to enforce linguistic consistency in high-stakes applications (like legal or educational software).

)4. Meta-Evaluation and Governance Framework (Addressing RQ1 & Broader Implications)

The system requires transparent governance tools based on the paper's findings.

Implement a continuous monitoring dashboard that tracks dialectal performance across different LLMs and domains, using the metrics derived from this study as benchmarks for fairness.

Specific Actions:

  1. Establish Domain-Specific Fairness Benchmarks: Create standardized test sets (e.g., formal NQ vs. informal ELI5) specifically designed to probe AmE/BrE performance gaps across various model architectures (GPT-4o, Llama, Claude).
  1. Develop a Hegemony Risk Alert System: Flag instances where the model’s generated text consistently exhibits strong AmE preference (>80% PAmE) in formal domains or when prompted with BrE, alerting researchers to potential linguistic homogenization risks.

Capability of Improved System:

This allows for proactive AI governance. Developers can identify which models are most susceptible to dialectal bias and prioritize fine-tuning efforts on those specific models or data sources, ensuring that the deployment of foundation models is more equitable and less prone to perpetuating sociopolitical biases encoded in colonial English hegemony.

Abstract

Large language models (LLMs) are increasingly embedded in educational, professional, and public infrastructure, yet widely used platforms expose "English (US)" as a primary English setting despite the global diversity of English. We ask: How does "English (US)" become the default? We study this question as structural bias, examining how geopolitical histories of data curation, digital dominance, and linguistic standardization intersect with the LLM development pipeline. Using British English as a controlled reference, we construct a curated resource of 1,813 matched American English (AmE)--British English (BrE) variants and introduce DiAlign, a dynamic, training-free method for estimating regional alignment from distributional evidence. We triangulate the AmE preference across data exposure --> representation --> generation, jointly examining pretraining and post-training data, tokenizer behavior and provenance, model prediction cost, and generated language across developer countries, prompt conditions, domains and sources, linguistic categories, and registers. AmE is consistently favored across all six audited pretraining corpora and 21 post-training datasets, is generally represented more compactly by tokenizers, and receives lower prediction cost. It also remains the dominant generation default under neutral English prompting; British-English prompting shifts this preference toward BrE but does not consistently eliminate the AmE default. To our knowledge, this is the first rigorous pipeline-wide study of structural bias across major phases of LLM development. Our findings show that contemporary LLMs privilege AmE as the de facto norm, raising concerns about linguistic homogenization, epistemic injustice, and inequity in global AI deployment, while providing a rigorous basis for targeted component-level intervention.

Sources

Related papers