Entangled in Representations: Mechanistic Investigation of Cultural Biases in Large Language Models
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Entangled in Representations: Mechanistic Investigation of Cultural Biases in Large Language Models".
Jane: The paper was written by Haeun Yu, Arnav Arora, Seogyeong Jeong, Nadav Borenstein, Siddhesh Pawar et al. from University of Copenhagen and KAIST.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back to the show, everyone. Today we're digging into a paper with a title that really sets the stage: "Entangled in Representations: Mechanistic Investigation of Cultural Biases in Large Language Models." Jane, when you first saw that word "entangled," what went through your mind?
Jane: Oh, Tom, it's perfect, honestly. It suggests that cultural biases aren't just something the model says, but something woven into its very fabric. The authors, from the University of Copenhagen and KAIST, are basically saying that if we want to fix cultural bias, we can't just look at the output text. We have to look inside the model's brain, so to speak.
Tom: Exactly. And that's a huge shift. Most prior work just asked the model questions and graded its answers. That's what they call extrinsic evaluation. But this paper is all about intrinsic evaluation—peeking at the hidden layers to see how cultural knowledge is actually stored and mixed together.
Jane: Right. And they introduce a method called Culturescope to do exactly that. It's like a tiny camera they insert into the model's processing pipeline to take snapshots of what's happening at each layer.
Lu: If I can jump in here, Tom. What excites me as a researcher is that they're not just saying "the model is biased." They're showing *where* that bias lives and *how* it flows. The title's "entangled" really captures the idea that knowledge about, say, Greece isn't stored in a neat little box. It's tangled up with knowledge about Turkey, about the UK, about everything else.
Meng: And from an engineering standpoint, that's both fascinating and terrifying. Because if the knowledge is entangled, then you can't just delete a bad memory or patch a single neuron. You have to untangle the whole mess.
Jane: That's a great point, Meng. And the paper actually gives us a way to measure that mess. They call it the Cultural Flattening score. It's a way to quantify how much one culture's distinctive knowledge is being overwritten or conflated with another culture's.
Tom: So it's not just "the model is biased," it's "the model is flattening Greek culture into Turkish culture" or something like that. That's a much more precise diagnosis.
Lu: Precisely. And that precision is what we need if we're ever going to build models that genuinely respect cultural diversity. We need to know the specific failure modes, not just a general sense of unfairness.
Meng: I'm curious about the practical side. How do they even get the model to "show" them its internal knowledge? That sounds like a technical nightmare.
Jane: That's the clever part, Meng. They use a technique called activation patching. They basically take the hidden state from one part of the model and inject it into another part, using a special prompt that asks the model to list cultural concepts. It's like asking the model to describe a picture it's holding in its mind, even if it can't show you the picture directly.
Tom: And the results are pretty striking. They found that the model's internal knowledge space is asymmetric. For example, knowledge about China flows into South Korea, but not the other way around. That's cultural flattening in action.
Lu: It's a powerful visualization. You can literally see the direction of the bias. And it's not random. It follows geographic and cultural proximity, which makes sense given the training data.
Meng: So the model is basically learning that "nearby cultures are interchangeable." That's a pretty deep-seated bias to have.
Jane: It is. And that's why this paper is so important. It's not just a new evaluation metric. It's a new way of seeing the problem. And once you can see it, you can start to fix it.
Tom: And that's exactly where we're headed next. We're going to talk about what they actually found when they put these models to the test.
Summary: Tom: So, Jane, we've set the stage with the title and the core idea. Now let's get into the meat of the paper. What did they actually do, and what did they find?
Jane: Well, Tom, they ran their Culturescope method on three different open-source models—Llama-three point one, aya-expanse, and Qwen2 point 5. And they used two cultural datasets, BLEnD and CAMeL-two covering a bunch of countries.
Lu: And the key finding, which I think is really profound, is that the models show a clear Western-dominance bias. When they're unsure, they tend to fall back on answers that are typical of high-resource, Western cultures.
Meng: So it's not just that they get things wrong. They get things wrong in a very specific, biased way. They're not making random mistakes.
Jane: Exactly. And they designed a clever experiment to prove it. They created multiple-choice questions with "hard negatives." So instead of giving the model one right answer and three obviously wrong ones, they gave it three plausible answers from different cultures.
Tom: Right. So for a question about China, they might offer the correct Chinese answer, but also the correct Korean answer and the correct American answer. That forces the model to actually distinguish between cultures, not just pick the most common answer.
Lu: And when the models got those questions wrong, they didn't pick randomly. They consistently picked the answer from the high-resource culture. That's the Western-dominance bias showing up in the output.
Meng: But here's the twist. They also looked at the attention maps—the internal focus of the model—and found the same pattern. The model's attention was literally drawn to the Western option, even before it made a mistake.
Jane: That's the really cool part. It means the bias isn't just a surface-level artifact. It's deeply embedded in how the model processes information. The model is internally "leaning" toward the Western answer.
Lu: And that's what they mean by "entangled in representations." The bias is in the weights, in the activations, in the attention patterns. It's everywhere.
Meng: So what about the cultural flattening score? Did that show the same thing?
Jane: It showed something related but distinct. The CF score measures how much one culture's distinctive knowledge is being used to represent another culture. And they found asymmetric connections. For example, China's knowledge flows into South Korea, but not the reverse.
Lu: And that's not just a random pattern. It's correlated with resource levels and geographic proximity. The model is essentially saying, "These cultures are similar, so I'll just use one to stand in for the other."
Tom: But here's a fascinating wrinkle they found. For low-resource cultures, like Assam or Azerbaijan, the model showed *less* bias. It didn't flatten them as much.
Meng: Wait, that sounds like good news. The model is less biased against them?
Jane: Not exactly. The paper argues it's because the model has so little knowledge about those cultures in the first place. It's not that it's being fair. It's that it has nothing to flatten *with*. There's no knowledge to be biased about.
Lu: That's a crucial distinction. The absence of bias isn't the same as the presence of fairness. It's just an absence of knowledge. And that's a very different problem to solve.
Tom: So we have two problems: one is cultural flattening, where the model overgeneralizes, and the other is cultural erasure, where the model just doesn't know anything. And they're both bad.
Meng: And they probably need different solutions. For flattening, you need to debias the model. For erasure, you need to add more data.
Jane: Exactly. And that's a big takeaway from this paper. It's not a one-size-fits-all problem. We need to understand the specific failure mode before we can fix it.
Tom: And that's what we're going to dig into next—what the paper suggests we do about all this.
Improvements: Tom: Alright, so we've established that the models are biased in specific ways. But what does this paper actually propose we *do* about it? Jane, what's the path forward?
Jane: Well, Tom, the paper doesn't give us a magic fix. But it does give us a roadmap. The first step is diagnosis. And that's what Culturescope is for. It's a tool to identify exactly where and how cultural bias is embedded in a model.
Lu: And that's more valuable than it sounds. Right now, if you want to make a model more culturally aware, you're kind of flying blind. You might add more data, or fine-tune on a specific culture, but you don't know if it's actually working on the right thing.
Meng: Right. It's like trying to fix a car engine without being able to open the hood. Culturescope opens the hood and shows you exactly which part is misfiring.
Jane: And once you can see that, you can start to think about targeted interventions. If you know that the model is flattening China into South Korea, you can specifically work on separating those two knowledge spaces.
Lu: The paper also suggests that we need different strategies for different resource levels. For high-resource cultures, the problem is overgeneralization. You need to debias, to untangle the knowledge. For low-resource cultures, the problem is underrepresentation. You need to add knowledge, not just debias.
Tom: So it's not just "make the model more diverse." It's "make the model more diverse in the right way, for the right reasons."
Meng: And that's a much more nuanced and practical approach. It acknowledges that cultural bias isn't a single problem. It's a collection of related problems that need different solutions.
Lu: Exactly. And the paper's cultural flattening score is a step toward making that distinction measurable. You can track whether your intervention is actually reducing flattening, or just making the model more confused.
Jane: Another interesting implication is for evaluation. This paper shows that just looking at accuracy on cultural questions isn't enough. You need to look at *how* the model is wrong, not just *that* it's wrong.
Tom: That's a big deal for the whole field of cultural AI evaluation. We need to move beyond simple benchmarks and start looking at the internal mechanisms.
Meng: I'm also thinking about the practical engineering side. These interpretability techniques are powerful, but they're also computationally expensive. Is this something that can be used in a real-world development pipeline?
Jane: That's a fair question. The paper uses eight-billion parameter models, which are manageable. But scaling this to larger models, or to many more cultures, would be a challenge.
Lu: But I think the value is more in the research direction. It tells us *what* to look for. Once we know that attention patterns are a reliable indicator of bias, we can develop cheaper proxies for that.
Tom: So it's not necessarily a tool you'd run on every model in production. But it's a tool you'd use to understand your model, to validate your training data, and to guide your fine-tuning.
Meng: That makes sense. It's a diagnostic tool, not a runtime tool. Like an MRI for the model.
Jane: Exactly. And the paper also validates Culturescope itself by showing that the cultural knowledge it extracts actually helps the model answer questions better. So it's not just an interesting toy. It's genuinely capturing useful information.
Lu: That's a crucial validation. It means the knowledge they're extracting from the internal layers is real and relevant, not just noise.
Tom: So the path forward is clear: use tools like Culturescope to diagnose, use the CF score to measure, and then develop targeted interventions based on the specific failure mode. That's a much more sophisticated approach than we've had before.
Jane: And it's a foundation for future work. The paper is very clear that this is just the beginning. There's a lot more to explore.
Meng: I'm curious to see how this scales to other languages and cultures, and whether the findings hold up in more diverse settings.
Tom: And that's a perfect segue into our final segment, where we wrap up and look at the big picture.
Conclusion: Tom: So, we've covered a lot of ground on "Entangled in Representations: Mechanistic Investigation of Cultural Biases in Large Language Models." Let's bring it all together.
Jane: The core message, Tom, is that cultural bias in LLMs is not a surface problem. It's deeply embedded in the model's internal representations. And to understand it, we need to look inside, not just at the outputs.
Lu: And the paper gives us the tools to do that. Culturescope lets us probe the internal knowledge, and the Cultural Flattening score lets us measure how cultures are being conflated.
Meng: The findings are clear: models show a Western-dominance bias, and they flatten low-resource cultures into their neighbors. But the paper also shows that low-resource cultures are less affected, simply because the model knows less about them.
Jane: That's the key nuance. It's not that the model is fair to low-resource cultures. It's that it's ignorant of them. And that's a different problem that needs a different solution.
Tom: So the takeaway is that we need targeted interventions. For high-resource cultures, we need to untangle the knowledge. For low-resource cultures, we need to add knowledge.
Lu: And we need to evaluate our models not just on accuracy, but on the *patterns* of their mistakes. This paper provides a framework for doing exactly that.
Meng: From an engineering perspective, this is a valuable diagnostic tool. It's not something you'd run in production, but it's something you'd use during development to understand your model's cultural blind spots.
Jane: And that's the exciting part. This isn't just an academic exercise. It has real implications for building AI systems that are more respectful of cultural diversity.
Tom: Absolutely. And with that, we're going to wrap up our discussion of this paper. It's a thought-provoking piece of research that opens up a whole new avenue for understanding and mitigating cultural bias in AI.
Jane: Thanks for joining us, everyone. We'll be back soon with another paper to dig into.
Tom: Until then, keep questioning, keep exploring, and keep looking inside the black box. Goodbye, everyone.
Haeun Yu, Arnav Arora, Seogyeong Jeong, Nadav Borenstein, Siddhesh Pawar, Jisu Shin, Jiho Jin, Junho Myung, Alice Oh, Isabelle Augenstein
University of Copenhagen · KAIST
cs.CL, cs.AI
Submitted: 2026-08-21
Updated: 2026-08-24
Comments: 16 pages, 7 figures
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 63/100
The gist: This paper introduces Culturescope, the first mechanistic interpretability-based method to probe the internal representations of cultural knowledge in large language models (LLMs), moving beyond
Key concepts
- Cultural Bias / Entanglement
- This suggests that cultural knowledge within a Large Language Model is not stored neatly but is 'tangled up' with other cultures. The paper argues that biases are woven into the model's internal structure, meaning simple output grading is insufficient for understanding or fixing the problem.
- Culturescope
- This is a method used in the paper to peek inside the model's hidden layers. It acts like a 'tiny camera' inserted into the processing pipeline, allowing researchers to take snapshots of how cultural knowledge is stored and mixed within each layer of the model.
- Cultural Flattening Score (CF score)
- This metric quantifies how much one culture's unique knowledge is being overwritten or conflated with another culture's knowledge. It helps measure the degree to which a model generalizes or merges distinct cultural concepts into a single, less precise representation.
- Western-Dominance Bias
- The research found that when faced with ambiguous questions, the models tend to default to answers typical of high-resource, Western cultures. This bias is not random; it is deeply embedded in the model's internal processing and attention patterns.
Terminology
Summary
This paper introduces Culturescope, the first mechanistic interpretability-based method to probe the internal representations of cultural knowledge in large language models (LLMs), moving beyond extrinsic evaluation of model outputs to uncover how cultural biases are encoded within model parameters. The authors propose a cultural flattening score
(CF score) to quantify the degree to which a target culture's distinctive knowledge is represented in a source culture, revealing asymmetric patterns of overgeneralization. They also analyze attention contribution scores to trace how Western-dominance bias and cultural flattening emerge internally.
The method operates in three stages: (1) Inference, where an LLM answers an open-ended cultural question; (2) Scoping-in, where activation patching (based on Patchscope) elicits a comma-separated list of cultural knowledge from the hidden representation of the answer; and (3) Filtering, where semantic similarity scores remove non-cultural knowledge. The CF score is computed using chi-square contributions to identify culturally distinctive knowledge and then summing contributions over overlapping knowledge between target and source cultures.
Experiments are conducted on two datasets—BLEnD (cultural commonsense QA) and CAMeL-2 (extractive QA)—across three open-source LLMs (Llama-3.1-8B-Instruct, aya-expanse-8b, Qwen2.5-7B-Instruct), covering English, Spanish, and Arabic. The authors create multiple-choice questions with hard negatives (BLEnD-Resource and BLEnD-Region) to simulate cultural biases, sampling options from different resource levels or geographical regions.
Key findings include: (1) CF score results show asymmetric connections among geographically or culturally proximate cultures (e.g., China → South Korea, Iran → Azerbaijan), indicating uneven generalization of culturally distinctive knowledge; (2) Attention contribution scores reveal Western-dominance bias in BLEnD-Resource and cultural flattening in BLEnD-Region, with statistically significant differences; (3) Extrinsic evaluations show that LLMs prefer biased options when answering incorrectly, but this susceptibility is weaker for low-resource cultures—likely due to limited parametric knowledge rather than improved fairness.
The paper validates Culturescope through relevance evaluation (prepending elicited knowledge improves accuracy over baselines like Cultural Prompting and CANDLE) and irrelevant patching (random noise vectors produce lower similarity scores). The authors conclude that low-resource cultures appear less susceptible to cultural biases because of insufficient representational coverage, suggesting future work should prioritize knowledge acquisition for these cultures rather than bias mitigation alone.
Improvements for AI systems
Based on the paper, here are the specific improvements I can implement in an AI system:
What I can do: Add a three-stage pipeline (Inference → Scoping-in → Filtering) that probes internal hidden representations rather than just decoding output text.
Specific implementation:
-
During inference, capture hidden states from all layers at the answer token positions
-
Weighted-sum the hidden states for noun/verb tokens to create a single representative vector
-
Patch this vector onto an inspection prompt's placeholder token to elicit comma-separated cultural knowledge
-
Filter generated knowledge using semantic similarity scores (cosine similarity between input text and generated concepts, keeping those above the mean)
Resulting capability: The system can now reveal why it gives a particular cultural answer, not just what it answers. For example, when asked about Greece's popular vacation spot, it can surface internal concepts like Rhodes,
Crete,
Aegean Sea
that it actually used in reasoning.
Abstract
The growing deployment of large language models (LLMs) across diverse cultural contexts necessitates a deeper understanding of LLMs' representations of different cultures. Prior work has focused on evaluating the cultural awareness of LLMs by only examining the text they generate. This approach overlooks the internal sources of cultural misrepresentation within the models themselves. To bridge this gap, we propose Culturescope, the first mechanistic interpretability-based method that probes the internal representations of different cultural knowledge in LLMs. We also introduce a cultural flattening score as a measure of the intrinsic cultural biases of the decoded knowledge from Culturescope. Additionally, we study how LLMs internalize cultural biases, which allows us to trace how cultural biases such as Western-dominance bias and cultural flattening emerge within LLMs. We find that low-resource cultures are less susceptible to cultural biases, likely due to the model's limited parametric knowledge. Our work provides a foundation for future research on mitigating cultural biases and enhancing LLMs' cultural understanding.
Sources
- Challenges and Strategies in Cross-Cultural NLP
- CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering