How Does "English (US)" Become the Default? Triangulating Structural Bias Towards American English Across the LLM Pipeline
summary
The gist
Large language models (LLMs) increasingly deploy only limited language settings, most notably “English (US),” despite global diversity and colonial history, raising foundational questions about
In short
This study investigated how large language models structurally prefer American English (AmE) despite global diversity. By auditing pretraining data, tokenizer designs, and model outputs using triangulation methods, researchers found a systematic bias toward AmE. This bias is deeply embedded in the training data and tokenization systems, leading LLMs to default to AmE even when prompted with British English.
Key concepts
- Pretraining Corpora Audits
- This involved systematically checking major datasets used to train LLMs for imbalances. Researchers quantified differences between American and British English by looking at word choices and alignment signals within these massive training sets. The results showed a consistent, statistically significant skew favoring American English.
- Tokenizers and Fertility
- Tokenizers break down text into smaller units (tokens). 'Fertility' measures how many subword tokens are needed to represent a word. The study found that British English forms require more tokens than American English, suggesting tokenization is less efficient for BrE, which impacts how the model processes different dialects.
- DIALIGN Method
- DIALIGN is a training-free method used to measure dialectal alignment. It analyzes n-grams (sequences of words) and calculates 'Signed Divergence' to predict which language variant a model will generate. This allowed researchers to quantify the generative preference for AmE versus BrE in model outputs.
- Generative Default
- This refers to the language that an LLM produces most often when given no specific instructions, or under a neutral condition. The study found that American English is the dominant generative default for these models, producing AmE outputs at rates between 65% and 80% with high confidence.
Terminology used across episodes
This episode discusses
- How Does "English (US)" Become the Default? Triangulating Structural Bias Towards American English Across the LLM Pipeline · Paper Radio
- Parity-Aware Byte-Pair Encoding: Improving Cross-lingual Fairness in Tokenization
- The Llama 3 Herd of Models · Paper Radio
- News Category Dataset
- OpinioRAG: Towards Generating User-Centric Opinion Highlights from Large-scale Online Reviews
- Beyond Fertility: Analyzing STRR as a Metric for Multilingual Tokenization Evaluation
- Llama 2: Open Foundation and Fine-Tuned Chat Models
The paper
How Does "English (US)" Become the Default? Triangulating Structural Bias Towards American English Across the LLM Pipeline · Read on arXiv
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "How Does "English (US)" Become the Default? Triangulating Structural Bias Towards American English Across the LLM Pipeline".
Jane: Large language models (LLMs) increasingly deploy only limited language settings, most notably “English (US),” despite global diversity and colonial history,
Tom: First, who's behind it and why it matters.
Title and authors: Tom: So, we're talking about "How Does 'English (US)' Become the Default? Triangulating Structural Bias Towards American English Across the LLM Pipeline," and it sounds like this paper is systematically checking every step of how an AI model learns its language. It’s not just one thing causing the bias; it’s a whole chain of decisions from data collection to final generation.
Jane: Exactly, Tom; what's striking is that they use a method called DIALIGN to estimate dialectal alignment using distributional evidence, which is designed to capture various contrasts like grammar and style simultaneously. It’s trying to get a holistic view of the bias rather than just looking at one area in isolation.
Lu: The authors are constructing this curated corpus of one thousand eight hundred thirteen AmE–BrE variants specifically for this investigation, which is a massive undertaking to capture that diversity before applying their triangulation method across the three stages they outline.
Meng: From an engineering standpoint, having them formalize the research questions around pretraining corpora audits and tokenizer representations gives us concrete checkpoints to test our own model architectures against these specific points of failure.
Lalam: I see this as a necessary postcolonial framing because it suggests that the way English is standardized and curated by digital dominance creates an epistemic injustice where one dialect, American English, gets implicitly prioritized.
The paper's summary: Tom: The core of the paper is that they are triangulating evidence across three main stages—pretraining corpora audits, tokenizer representations, and generative preferences—to show how structural bias toward American English is introduced and amplified throughout the LLM development pipeline.
Jane: They found that in the pretraining data stage, there's a systematic skew toward AmE, especially with orthographic variants like color versus colour where AmE spellings dominate by margins above seventy percent. That immediately tells us the foundation is already tilted.
Lu: And then they move to the tokenizer stage, revealing that British English forms have higher fertility than their American counterparts, which points to less efficient tokenization for BrE and a gap of about eighteen point seven two percent in vocabulary-based differences.
Meng: That fertility metric is very tangible; it means the tokenizer isn't handling certain British words as efficiently as others, which translates directly into computational overhead or potentially less nuanced representation for those forms.
Lalam: It’s a really important finding because it shows that the problem isn't just in the model itself, but in how we tokenize the language before it even enters the model's core structure.
The paper's improvements: Tom: The authors suggest some really practical steps for fixing this, focusing on three areas: building dialect-sensitive corpus construction, injecting BrE tokens into base tokenizers using DIALIGN, and being extremely careful about adopting pretrained tokenizers without regional adaptation.
Jane: It seems like the authors are pushing for a proactive approach rather than just analyzing the problem; they want us to actively design systems that respect regional linguistic differences from the start.
Lu: The suggestion about using DIALIGN to inject BrE tokens into base tokenizers is ambitious, but it proposes a way to dynamically expand the vocabulary based on distributional evidence, which is a creative use of that training-free method.
Meng: For us engineers, the recommendation against blindly adopting pretrained tokenizers without adaptation is crucial; it tells us we can't just plug and play a tokenizer and assume fairness across dialects.
Lalam: I think the most profound improvement they suggest relates to how we approach the system itself—they are advocating for component-wise design recommendations to prevent linguistic homogenization in the broader AI deployment landscape.
Conclusion: Tom: So, to wrap up, this paper on "How Does 'English (US)' Become the Default? Triangulating Structural Bias Towards American English Across the LLM Pipeline" shows that dialectal skew isn't accidental; it’s structurally embedded in the pipeline through data choice, tokenization inefficiencies for British English forms, and generative defaults.
Jane: It really underscores the risk of linguistic homogenization if we don't actively work on building more inclusive systems; we need to be aware of how these models might enforce specific sociolinguistic norms in high-stakes environments.
Lu: I think the long-term vision here is that by understanding this pipeline structure, we can move toward truly dialect-sensitive corpus construction and tokenization design that reflects global English diversity rather than just one dominant standard.
Meng: Practically speaking, the implication for us is a need to build in those audit layers we discussed earlier so we can flag when our current models are showing those high AmE preferences before they go into widespread use.
Lalam: Ultimately, this paper gives us a roadmap to prevent epistemic injustice by ensuring that the underlying AI infrastructure respects the full range of English varieties, making the cultural impact of these systems much fairer for everyone involved.
More episodes
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language
- 2508.08833-An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
- 2405.04118-Policy Learning with a Language Bottleneck