Who Brought Easter Eggs to Eid? Auditing LLM-Generated Cultural Translation of Math Word Problems Across Languages and Regions
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Who Brought Easter Eggs to Eid? Auditing LLM-Generated Cultural Translation of Math Word Problems Across Languages and Regions".
Jane: Large language models are increasingly used to adapt math word problems for personalized learning at scale, but it remains an open question whether those adaptations are consistent across models,
Tom: First, who's behind it and why it matters.
Title and authors: Tom: So we've established what this paper is about—it’s an audit of how LLMs adapt math word problems across different languages and regions. The title itself, "Who Brought Easter Eggs to Eid?", really captures the kind of subtle, sometimes off-base, cultural substitutions they found happening. The authors are Parisa Suchdev, Juniper Lovato, and others from the Computational Ethics Lab at the University of Vermont.
Jane: Yes, and when we look at the title in that way, it immediately signals that they aren't just doing simple word-for-word translation; they’re looking for those moments where cultural knowledge gets mixed up with mathematical logic. It sets a really interesting tone for what they are testing.
Lu: The authors are from a strong computational science background, which is perfect because they clearly have the technical chops to design such a complex audit across three different frontier models simultaneously. That kind of cross-model comparison is exactly what we need to see more of.
Meng: I wonder if their specific choice of entities—names, foods, places—is intentional in how they structured the sixty problems they used from GSM8K <ref:2606.11009#pg0>? That selection process could really dictate which cultural biases surface in the results.
Lalam: I think it’s important to remember that these authors are looking at how models treat cultural entities as something distinct from pure mathematical notation, which is a fundamental distinction for any AI aiming to be truly helpful across cultures.
The paper's summary: Tom: Moving on, the core of this study involves auditing how Claude Opus four GPT four point one, and Gemini two point five Pro handle adapting those math problems into seven different target languages <ref:2606.11009#pg0,how Claude Opus 4, GPT 4.1, and Gemini 2.5 Pro>. Essentially, they took sixty English word problems and had the models localize them based on a prompt asking them to fit the cultural context of students in specific countries like India or Italy.
Jane: What I find interesting is that the researchers are not just checking if the output is grammatically correct; they are meticulously coding whether the model preserves, localizes, generalizes, or omits entities like names and foods during this process. It’s a very detailed look at the decision-making layer of an LLM.
Lu: The paper highlights that adaptation isn't just one step but a sequence of choices—preserving something universal versus localizing it to fit the specific region. This suggests that cultural context influences how well these models actually solve math problems, which is a pretty deep finding because it ties language directly into mathematical performance.
Meng: I see the focus on the results where models agree on transformation type in sixty-two point five percent of cases and specific substitutions in thirty-three point five percent, which points to a lot of inconsistency across different model architectures. That variability is something engineers have to worry about when trying to build reliable systems.
Lalam: That lack of agreement is actually key because it shows that the choice between models directly shapes the cultural world a student encounters, as the paper suggests. It’s not one universal translation rule; it’s model-specific adaptation strategy.
The paper's improvements: Tom: The researchers point out some things they think could make this audit even better, focusing on how we can improve these models' performance in this area. They suggest building an advanced cultural consistency layer that goes beyond simple checks to ensure models agree on the specific choices they make across different runs.
Jane: That idea of a Cultural Output Validator module sounds really practical for improving reliability. It’s about forcing the AI to stick to a consistent set of local entities once it starts adapting, rather than letting it drift randomly between iterations.
Lu: I like that direction because it moves the focus from just observing what happened to actively designing systems that enforce desired cultural constraints in the adaptation pipeline. That kind of active control over localization is where things get really interesting for future AI.
Meng: From a practical implementation angle, building something that can reliably enforce agreement on these substitutions across different linguistic outputs sounds like a big undertaking, but if it works, it could significantly reduce the need for manual cultural review down the line.
Lalam: I think the core improvement they suggest is moving toward an entropy-aware diversity control mechanism. If we want AI to be good at cultural representation, we have to actively stop it from collapsing all that variety into just a few common choices.
Conclusion: Tom: So, wrapping up this discussion on "Who Brought Easter Eggs to Eid? Auditing LLM-Generated Cultural Translation of Math Word Problems Across Languages and Regions," the paper confirms that model choice really does dictate which cultural world students see when their math problems are adapted. They show that models tend to prioritize certain entities like personal names, even if they struggle with broader cultural practices.
Jane: Exactly, and the main implication is that we need to be much more careful about surface markers; what looks locally appropriate on the surface might hide deeper issues regarding authentic cultural understanding. We have to scrutinize these adaptations closely because models can use "surface plausibility" to mask real failures in understanding context.
Lu: The paper strongly suggests that as we move toward personalized learning at scale, we can’t just rely on a single model; we need systems that are designed with explicit controls over cultural attention and diversity budgets to ensure they don't accidentally homogenize the student experience.
Meng: For me, the practical implication is clear: when deploying these tools, we have to build in mechanisms for structured pedagogical feedback. We shouldn't just take the output; we need a rationale explaining *why* an entity was preserved or changed so educators can actually make sense of it.
Lalam: Ultimately, this work shows us that AI can be a powerful tool for personalization only if we treat cultural localization as a complex sequence of deliberate choices rather than an automatic translation. The findings from this paper are crucial for building systems that respect and retain the richness of human culture across all languages.
Computational Ethics Lab · Department of Computer Science, University of Vermont · Vermont Complex Systems Center, University of Vermont
cs.CL, cs.CY
Submitted: 2026-06-09
Updated: 2026-10-07
Comments: 18 pages total with references and appendix, 9 figures, accepted at AIES
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 82/100
The gist: Large language models are increasingly used to adapt math word problems for personalized learning at scale, but it remains an open question whether those adaptations are consistent across models,
Key concepts
- Cultural Localization
- This is the process where an AI adapts a problem by replacing English-specific details (like names or currencies) with equivalents relevant to a target culture. The goal is to make the math problem feel authentic and understandable to students in a specific country, rather than just translating words.
- Entropy Collapse
- This measures how much variety is lost when translating problems across different languages. The study found that as problems are adapted, the range of cultural entities present in the original English set shrinks. This means that instead of seeing many different cultural concepts, students encounter a smaller, more similar set of substitutes.
- Entity Salience
- This refers to which specific elements—like names or foods—the LLMs prioritize when making changes. The study found that personal names and culinary practices are treated as overwhelmingly important by the models, suggesting these surface markers carry the most weight in their cultural adaptations.
Terminology
Summary
Large language models are increasingly used to adapt math word problems for personalized learning at scale, but it remains an open question whether those adaptations are consistent across models, preserve cultural diversity at scale, and reveal which cultural entities models treat as most salient. This study analyzes how three frontier LLMs—Claude Opus 4, GPT 4.1, and Gemini 2.5 Pro—adapt English math word problems into seven diverse languages across various regions to determine if model choice dictates the cultural world students encounter.
How it works
The researchers audited how three frontier LLMs adapt 60 English GSM8K math word problems into seven target languages: Bengali, Hindi, and Punjabi (India); Urdu and Sindhi (Pakistan); and Italian and Sicilian (Italy). The process involved generating 1,260 translated problems using a teacher-centric prompt instructing the models to Translate the following math word problem from English into [language] and adapt the problem so that it fits the cultural context of students in [country].
The goal was not just translation but cultural localization, involving substitutions for names, foods, currencies, places, institutions, and practices.
The core of the analysis is an entity-level dataset comprising 6,489 transformations. For each English source problem from GSM8K (a dataset of 8.5K high-quality grade school math word problems), researchers identified culturally adaptable entities—including names, foods, currencies, places, institutions, and practices—and aligned them with their target-language realizations. This mapping was performed using an LLM-assisted procedure with GPT-4.1 to narrow the search space before manual validation ensured that every source key appeared exactly once in the target output.
Research Questions Addressed
The study was structured around three primary research questions:
-
When the same mathematics word problem is given to multiple frontier LLMs with the same instruction, do they produce the same cultural output? This investigates whether model choice is a
cultural decision.
-
When English math word problems are adapted into another language, do the outputs preserve the range of cultural entities present in the source set, or does adaptation paradoxically compress that variety into a narrower and more homogeneous target world? This measures
entropy collapse
between source and translated entities. -
Which entities do LLMs appear to treat as carrying the most cultural weight? This examines how model-internal priorities resemble those of a teacher when making pedagogical judgments about which elements to preserve, localize, generalize, or omit.
Key Findings on Model Behavior
The analysis revealed significant discrepancies in model behavior across several dimensions. Models agree on transformation type in only 62.5% of cases and on specific substitutions in only 33.5%, indicating that model choice directly shapes which cultural world students encounter.
Cross-model agreement on both action and value is only found in 33.5% of instances, with pairwise agreement rates varying significantly (e.g., Claude and GPT are most similar, while Gemini and GPT are least similar).
The study identified systematic patterns:
-Model Tendencies:
Claude Opus 4 localizes most aggressively across languages,
while GPT-4.1 preserves most conservatively,
and Gemini 2.5 Pro shows the highest type-changed rates.
-Entity Salience:
Personal Names showed the strongest reduction in diversity, with a negative entropy difference of-2.638 bits across nearly all models and languages, suggesting they are treated as overwhelmingly salient. Other categories showing significant collapse include Life Course Milestones (-0.593 bits) and Culinary Practices (-0.45 bits).
-Cultural Contamination:
Surface plausibility masks deeper failures; for instance, models adapted an egg hunt
as an Eid activity,
demonstrating pattern-matching without cultural knowledge of accompanying practices, which the researchers noted is not authentic.
Implications for Pedagogy and Research
The findings suggest that LLMs prioritize surface markers—names, foods, and currencies—while preserving deeper structural features like grade-level systems that embed culturally specific assumptions. The results show that adaptation systematically reduces cultural diversity across all 21 language-model combinations (negative entropy differences). While models often change entity values to make problems seem locally appropriate, they repeatedly draw from a limited set of substitutes,
leading to outputs that appear locally appropriate in individual examples but become more homogeneous overall. This necessitates active cultural review by educators, as the surface plausibility is precisely what makes deeper failures easy to overlook. The study concludes that the systems educators rely on often hide instability and cultural compression beneath a surface of local appropriateness, making scrutiny a necessary step in the adaptation process.
Limitations
The study's limitations include operating at the entity level rather than analyzing narrative coherence, meaning it captures individual transformations but not the broader educational impact.
Improvements for AI systems
As a fastidious and diligent researcher, I have analyzed this paper, Who Brought Easter Eggs to Eid? Auditing Cultural Translation of Math Word Problems Across Diverse Languages and Regions,
focusing on its empirical findings regarding Large Language Models (LLMs) in culturally localized content adaptation.
Based on the research questions (RQ1: Consistency; RQ2: Variety/Entropy Collapse; RQ3: Salience), here are specific, actionable improvements for AI systems and what those improved systems can achieve.
)
- Advanced Cultural Consistency Layer (Addressing RQ1)
-
The system should implement a
Cultural Output Validator
module that goes beyond simple grammaticality checks. This module would be trained on the entity-level dataset (6,489 transformations). -
It must enforce agreement on both the action taken (Preserve, Localize, etc.) and the specific substitution value across multiple model calls for the same source problem.
-
Improved AI System Capability: The system will produce
Culturally Consistent Adaptations,
ensuring that if a user requests an adaptation for Sicilian speakers, the resulting problem uses a consistent set of local entities (e.g., always usingeuro
instead of sometimes defaulting totaka
), rather than allowing models to drift between iterations.
- Entropy-Aware Diversity Control Mechanism (Addressing RQ2)
-
Implement a dynamic diversity metric based on Shannon entropy applied to the entity value distribution across the target corpus, similar to the study's analysis.
-
The system should be equipped with a
Diversity Budget
parameter. If an adaptation process shows signs of entropy collapse (negative change), it should trigger an internal mechanism to intentionally select less-frequent, but culturally relevant, substitutions from a curated knowledge base instead of defaulting to the highest probability (most frequent) substitution. -
Improved AI System Capability: The system will prevent
cultural homogenization.
It will actively resist the tendency of LLMs to repeatedly select the same names or foods across different tasks, ensuring that adaptations for diverse linguistic groups retain a broader and more representative range of cultural entities.
- Salience-Weighted Cultural Attention Prioritization (Addressing RQ3)
-
Develop an internal
Cultural Weight Matrix
derived from the entity-level analysis (e.g., Personal Names show high action agreement but low value agreement, suggesting models prioritize the type/action over specific value). -
The system should allow for explicit weighting of entity categories based on pedagogical goals (e.g., if the goal is to teach currency, weight "Money & Currency Systems
higher; if the goal is cultural immersion, weight
Food & Drink Items"). -
Improved AI System Capability: The system will allow educators/users to specify the desired cultural focus. If a teacher wants a problem focused on Italian culture, the system can be fine-tuned to prioritize localization for entities in that category (like
restaurant
orapparel
) while allowing other categories likevehicle parts
to remain preserved, mimicking the contested middle observed in Table 4.
- Contextual Misattribution Guardrails (Addressing Surface Plausibility)
-
Integrate a regional context verification step into the prompt execution pipeline. This step would cross-reference the specified target country/language with known cultural data biases (e.g., checking if a specified Indian language adaptation defaults to Bangladeshi currency).
-
The system should employ
Cross-Cultural Contamination Checks
that flag adaptations where surface markers (like names or holidays) are used without ensuring the underlying practice is culturally authentic to the target setting. -
Improved AI System Capability: The system will mitigate
surface plausibility errors.
It will actively check for and correct regional misattributions, ensuring that an adaptation intended for students in India does not inadvertently use Bangladeshi context markers (like Taka) or substitute Western holidays with inappropriate local equivalents (like adapting an egg hunt to Eid activities).
- Structured Pedagogical Feedback Loop
-
The output should not just be the adapted problem, but a
Cultural Rationale Report.
This report would explicitly state:We localized Entity X using Value Y because it is culturally salient in Region Z,
orEntity A was preserved because it relates to a universal concept (e.g., grade level).
-
Improved AI System Capability: The system moves from being a mere generator to an informed pedagogical assistant. It provides the user with transparency into the model's cultural judgment, allowing the educator to make informed, critical decisions about whether the model's choices reflect genuine cultural grounding or merely shallow substitution.
Abstract
Large language models are increasingly used to adapt math word problems for personalized learning at scale, but it remains an open question whether those adaptations are consistent across models, preserve cultural diversity at scale, and reveal which cultural entities models treat as most salient. We analyze how Claude Opus 4, GPT-4.1, and Gemini 2.5 Pro adapt 60 English math word problems into Bengali, Hindi, Punjabi (India), Urdu, Sindhi (Pakistan), Italian, and Sicilian (Italy), a language set spanning the full resource spectrum, from high-resource Italian and Hindi to under-studied Sindhi, Sicilian, and Punjabi. We annotate 6,489 entity transformations, coding whether models preserve, localize, generalize, omit, or change entities such as names, foods, and places. Models agree on transformation type in 62.5% of cases and on specific substitutions in only 33.5%, meaning model choice directly shapes which cultural world students encounter. All 21 language-model combinations show entropy collapse, with adaptation compressing rather than expanding cultural diversity. Models prioritize surface markers such as names, foods, and currencies while preserving deeper structural features such as grade-level systems that embed culturally specific assumptions. Despite prompts specifying target countries, models misattribute regional context by using Bangladeshi taka for Indian Bengali students and produce cross-cultural contamination, such as adapting egg hunts as Eid activities. Some failures are visible in individual translations. Others, including diversity collapse, systematic preference for surface markers, and consistent regional misattribution, emerge only through corpus-level analysis. The surface plausibility that makes adapted problems look correct is precisely what makes deeper failures easy to overlook.
Sources
- Bridging the Culture Gap: A Framework for LLM-Driven Socio-Cultural Localization of Math Word Problems in Low-Resource Languages
- Training Verifiers to Solve Math Word Problems
- Towards Measuring the Representation of Subjective Global Opinions in Language Models
- Lost in Cultural Translation: Do LLMs Struggle with Math Across Cultural Contexts?
- CULTURE-GEN: Revealing Global Cultural Perception in Language Models through Natural Language Prompting
- Mathematics Isn't Culture-Free: Probing Cultural Gaps via Entity and Scenario Perturbations
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering