Who Brought Easter Eggs to Eid? Auditing LLM-Generated Cultural Translation of Math Word Problems Across Languages and Regions

summary

Video file (mp4)

The gist

Large language models are increasingly used to adapt math word problems for personalized learning at scale, but it remains an open question whether those adaptations are consistent across models,

In short

Researchers tested how three different large language models adapt English math word problems into seven diverse languages for students in various regions. The study found that model choice significantly impacts cultural output, and LLMs tend to treat personal names and foods as the most culturally important elements, often compressing cultural variety into a narrower set of substitutions.

Key concepts

Cultural Localization
This is the process where an AI adapts a problem by replacing English-specific details (like names or currencies) with equivalents relevant to a target culture. The goal is to make the math problem feel authentic and understandable to students in a specific country, rather than just translating words.
Entropy Collapse
This measures how much variety is lost when translating problems across different languages. The study found that as problems are adapted, the range of cultural entities present in the original English set shrinks. This means that instead of seeing many different cultural concepts, students encounter a smaller, more similar set of substitutes.
Entity Salience
This refers to which specific elements—like names or foods—the LLMs prioritize when making changes. The study found that personal names and culinary practices are treated as overwhelmingly important by the models, suggesting these surface markers carry the most weight in their cultural adaptations.

Terminology used across episodes

This episode discusses

The paper

Who Brought Easter Eggs to Eid? Auditing LLM-Generated Cultural Translation of Math Word Problems Across Languages and Regions · Read on arXiv

Computational Ethics Lab · Department of Computer Science, University of Vermont · Vermont Complex Systems Center, University of Vermont

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Who Brought Easter Eggs to Eid? Auditing LLM-Generated Cultural Translation of Math Word Problems Across Languages and Regions".

Jane: Large language models are increasingly used to adapt math word problems for personalized learning at scale, but it remains an open question whether those adaptations are consistent across models,

Tom: First, who's behind it and why it matters.

Title and authors: Tom: So we've established what this paper is about—it’s an audit of how LLMs adapt math word problems across different languages and regions. The title itself, "Who Brought Easter Eggs to Eid?", really captures the kind of subtle, sometimes off-base, cultural substitutions they found happening. The authors are Parisa Suchdev, Juniper Lovato, and others from the Computational Ethics Lab at the University of Vermont.

Jane: Yes, and when we look at the title in that way, it immediately signals that they aren't just doing simple word-for-word translation; they’re looking for those moments where cultural knowledge gets mixed up with mathematical logic. It sets a really interesting tone for what they are testing.

Lu: The authors are from a strong computational science background, which is perfect because they clearly have the technical chops to design such a complex audit across three different frontier models simultaneously. That kind of cross-model comparison is exactly what we need to see more of.

Meng: I wonder if their specific choice of entities—names, foods, places—is intentional in how they structured the sixty problems they used from GSM8K <ref:2606.11009#pg0>? That selection process could really dictate which cultural biases surface in the results.

Lalam: I think it’s important to remember that these authors are looking at how models treat cultural entities as something distinct from pure mathematical notation, which is a fundamental distinction for any AI aiming to be truly helpful across cultures.

The paper's summary: Tom: Moving on, the core of this study involves auditing how Claude Opus four GPT four point one, and Gemini two point five Pro handle adapting those math problems into seven different target languages <ref:2606.11009#pg0,how Claude Opus 4, GPT 4.1, and Gemini 2.5 Pro>. Essentially, they took sixty English word problems and had the models localize them based on a prompt asking them to fit the cultural context of students in specific countries like India or Italy.

Jane: What I find interesting is that the researchers are not just checking if the output is grammatically correct; they are meticulously coding whether the model preserves, localizes, generalizes, or omits entities like names and foods during this process. It’s a very detailed look at the decision-making layer of an LLM.

Lu: The paper highlights that adaptation isn't just one step but a sequence of choices—preserving something universal versus localizing it to fit the specific region. This suggests that cultural context influences how well these models actually solve math problems, which is a pretty deep finding because it ties language directly into mathematical performance.

Meng: I see the focus on the results where models agree on transformation type in sixty-two point five percent of cases and specific substitutions in thirty-three point five percent, which points to a lot of inconsistency across different model architectures. That variability is something engineers have to worry about when trying to build reliable systems.

Lalam: That lack of agreement is actually key because it shows that the choice between models directly shapes the cultural world a student encounters, as the paper suggests. It’s not one universal translation rule; it’s model-specific adaptation strategy.

The paper's improvements: Tom: The researchers point out some things they think could make this audit even better, focusing on how we can improve these models' performance in this area. They suggest building an advanced cultural consistency layer that goes beyond simple checks to ensure models agree on the specific choices they make across different runs.

Jane: That idea of a Cultural Output Validator module sounds really practical for improving reliability. It’s about forcing the AI to stick to a consistent set of local entities once it starts adapting, rather than letting it drift randomly between iterations.

Lu: I like that direction because it moves the focus from just observing what happened to actively designing systems that enforce desired cultural constraints in the adaptation pipeline. That kind of active control over localization is where things get really interesting for future AI.

Meng: From a practical implementation angle, building something that can reliably enforce agreement on these substitutions across different linguistic outputs sounds like a big undertaking, but if it works, it could significantly reduce the need for manual cultural review down the line.

Lalam: I think the core improvement they suggest is moving toward an entropy-aware diversity control mechanism. If we want AI to be good at cultural representation, we have to actively stop it from collapsing all that variety into just a few common choices.

Conclusion: Tom: So, wrapping up this discussion on "Who Brought Easter Eggs to Eid? Auditing LLM-Generated Cultural Translation of Math Word Problems Across Languages and Regions," the paper confirms that model choice really does dictate which cultural world students see when their math problems are adapted. They show that models tend to prioritize certain entities like personal names, even if they struggle with broader cultural practices.

Jane: Exactly, and the main implication is that we need to be much more careful about surface markers; what looks locally appropriate on the surface might hide deeper issues regarding authentic cultural understanding. We have to scrutinize these adaptations closely because models can use "surface plausibility" to mask real failures in understanding context.

Lu: The paper strongly suggests that as we move toward personalized learning at scale, we can’t just rely on a single model; we need systems that are designed with explicit controls over cultural attention and diversity budgets to ensure they don't accidentally homogenize the student experience.

Meng: For me, the practical implication is clear: when deploying these tools, we have to build in mechanisms for structured pedagogical feedback. We shouldn't just take the output; we need a rationale explaining *why* an entity was preserved or changed so educators can actually make sense of it.

Lalam: Ultimately, this work shows us that AI can be a powerful tool for personalization only if we treat cultural localization as a complex sequence of deliberate choices rather than an automatic translation. The findings from this paper are crucial for building systems that respect and retain the richness of human culture across all languages.

More episodes

← Home