Structural Silence: When AI Infrastructure Fails Speakers of Underrepresented Languages
Avijit Roy, Proma Roy
John Jay College of Criminal Justice, CUNY · The City College of New York, CUNY
cs.CL, cs.AI, cs.CY
Submitted: 2026-08-12
Updated: 2026-08-13
Comments: An associated poster version of this work was presented at the 69th Annual Conference of the International Linguistic Association (ILA 2026), New York, NY, April 30-May 2, 2026
Project page: https://w3techs.com/technologies/overview/content_language
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 75/100
The gist: This paper examines the structural barriers that disadvantage speakers of underrepresented languages in AI infrastructure, using Bengali as a case study.
Terminology
Summary
This paper examines the structural barriers that disadvantage speakers of underrepresented languages in AI infrastructure, using Bengali as a case study. The authors identify four interlocking structural failures
that systematically exclude Bengali from AI systems before any model is trained:
-
The Web Presence Gap: Bengali speakers represent roughly 4% of the global population—approximately 242 million native speakers—yet Bengali accounts for fewer than 0.5% of global web content, while English accounts for approximately 49.5% of global web content despite having a comparable share of native speakers.
-
The Training Token Deficit: The Sangraha corpus allocates approximately 30 billion tokens to Bengali, while large English corpora such as Common Corpus reach approximately 2 trillion English tokens, creating a 67:1 ratio between English and Bengali token availability in major training resources.
-
The Tokenization Penalty: Standard tokenizers (BPE and WordPiece) were designed for Latin-script languages, but Bengali employs an alphasyllabary script with diacritics and conjunct forms that do not decompose cleanly under standard tokenization schemes. This results in elevated token fertility—Bengali text requires significantly more subword tokens to represent the same semantic content as English—which increases computational overhead and disrupts linguistic units models rely on.
-
The Connectivity Exclusion: Cloud-dependent AI tools are inaccessible to rural populations most likely to benefit from them. Individual internet penetration in rural Bangladesh stands at 36.5% compared to 71.4% in urban areas, with only 47.2% of the population counted as internet users by the end of 2024. Only 9.2% of households own a computer, and supplementary duty on mobile data increased from 3% in FY2016 to 23% in 2024.
The paper argues these failures "are not incidental byproducts of early-stage technology. They are structural features of the training corpora that models learn from, the tokenizers that process Bengali text, the benchmarks used to evaluate model performance, and the deployment architectures that assume reliable internet connectivity. Together, they constitute what the authors call
structural silence"—the systematic exclusion of a language from AI infrastructure not through explicit policy but through the accumulated weight of design decisions that were never made with that language in mind.
The educational consequences are framed through Cognitive Load Theory: when AI tools deliver explanations in English, learners with limited English proficiency must process both linguistic and technical content concurrently, which frequently exceeds working memory capacity, reducing both retention and the learner's ability to apply new knowledge to novel problems.
The authors cite Roussel et al. (2017), who demonstrated that content presented in a foreign language produced significantly lower outcomes across both content and language learning compared to native-language presentation.
The paper concludes with three responses: (1) low-resource language infrastructure work—datasets, benchmarks, and evaluation protocols—warrants recognition as primary research, not supporting labor; (2) offline-capable design should be treated as a serious educational access strategy with its own design requirements and evaluation criteria; (3) linguistic analysis should play a central role in identifying tokenization, evaluation, and modeling assumptions that remain invisible when English is treated as the reference standard. The authors emphasize that performance differences often reflect differences in training feasibility rather than differences in linguistic complexity or user demand,
and addressing inequity will require sustained institutional commitment to expanding dataset development for underrepresented languages and supporting infrastructure that makes large-scale training feasible.
Improvements for AI systems
Improvements to AI Systems:
- Language-Aware Tokenizer Selection and Training
-
Implement tokenizers specifically designed for alphasyllabary scripts (e.g., Bengali, Hindi, Tamil) using grapheme-cluster or syllable-based segmentation, rather than BPE/WordPiece optimized for Latin scripts.
-
Train tokenizers on native-language corpora to reduce token fertility by at least 30–50% for Bengali, lowering computational cost and preserving linguistic units (e.g., conjuncts, diacritics).
-
Resulting capability: AI systems can process Bengali text with fewer tokens, faster inference, and more accurate semantic representation, enabling better performance in translation, summarization, and question-answering.
- Offline-First Model Architecture and Deployment
-
Develop lightweight, quantized models (e.g., sub-1GB) that run entirely on-device, with no cloud dependency, using edge-optimized inference engines.
-
Pre-package offline language packs for underrepresented languages, including Bengali, with local retrieval-augmented generation (RAG) over a small, curated knowledge base.
-
Resulting capability: Rural users with limited or no internet can access AI tutoring, text-to-speech, and language learning tools in Bengali, improving educational outcomes without connectivity constraints.
- Equity-Aware Benchmarking and Evaluation
-
Create and integrate benchmarks that measure performance on low-resource languages using native-speaker-validated tasks (e.g., reading comprehension, grammar correction, and code-switched dialogue) rather than translated English tasks.
-
Report model performance separately for each language, with metrics adjusted for tokenizer efficiency and training corpus size, to expose structural gaps.
-
Resulting capability: AI developers can identify and prioritize improvements for underrepresented languages, leading to models that are genuinely multilingual rather than English-centric.
- Bilingual Cognitive Load Reduction in AI Interfaces
-
Design AI interfaces that automatically detect user proficiency and switch to native-language explanations, with optional code-switching for technical terms.
-
Use adaptive pacing and chunking of content to reduce working memory overload, based on cognitive load theory, for learners with low English proficiency.
-
Resulting capability: AI tutoring systems can deliver STEM or vocational content in Bengali with minimal cognitive interference, improving retention and problem-solving skills for non-English speakers.
- Structural Silence Auditing Tool
-
Build an automated audit framework that scans training corpora, tokenizers, and deployment pipelines for biases against underrepresented languages (e.g., token fertility ratios, web content share, connectivity assumptions).
-
Generate actionable reports with specific fixes (e.g.,
increase Bengali corpus by X tokens
orswitch to syllable-based tokenization
). -
Resulting capability: AI organizations can systematically identify and correct exclusionary practices before model release, ensuring fairer access across languages.
Abstract
Artificial intelligence tools for education and language support are increasingly framed as scalable responses to access gaps in under-resourced communities. Yet the infrastructure underlying these tools, including training corpora, tokenization schemes, evaluation benchmarks, and deployment architectures, can systematically disadvantage speakers of underrepresented languages before a model is trained. This paper examines these structural barriers through Bengali, one of the world's most widely spoken languages, focusing on AI-assisted education in low-connectivity environments. We identify four interlocking failures: a severe web presence gap, with Bengali accounting for less than 0.5% of global web content despite representing nearly 4% of the global population; a 67:1 training-token deficit between English and Bengali in major multilingual corpora; a tokenization penalty associated with Bengali's alphasyllabary script that compounds the data deficit through higher token fertility; and connectivity exclusion, with individual internet penetration at 36.5% in rural areas compared with 71.4% in urban areas. These failures reflect longstanding resource-allocation decisions, institutional priorities, and design defaults that did not center underrepresented languages in mainstream AI development. We argue that dataset scarcity should be understood as a structural barrier rather than an isolated technical limitation, and that offline-first design should be treated as an equity-oriented infrastructure strategy. We conclude with directions for linguistics and AI research aimed at reducing these structural inequalities.
Sources
- Script Fragmentation and Format: What Drives the English-Bengali Performance Gap in Open LLMs?
- Goldfish: Monolingual Language Models for 350 Languages
- QLoRA: Efficient Finetuning of Quantized LLMs
- The Teacher's Dilemma: Balancing Trade-Offs in Programming Education for Emergent Bilingual Students
- LoRA: Low-Rank Adaptation of Large Language Models
- Common Corpus: The Largest Collection of Ethical Data for LLM Pre-Training
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering