Entangled in Representations: Mechanistic Investigation of Cultural Biases in Large Language Models
summary
The gist
This paper introduces Culturescope, the first mechanistic interpretability-based method to probe the internal representations of cultural knowledge in large language models (LLMs), moving beyond
In short
The episode reviews a paper investigating how cultural biases are deeply embedded within Large Language Models (LLMs). The hosts discuss tools like Culturescope, which reveals that models exhibit Western-dominance bias and conflate cultures. The conclusion is that fixing these issues requires targeted interventions: untangling knowledge for high-resource cultures and adding data for low-resource ones.
Key concepts
- Cultural Bias / Entanglement
- This suggests that cultural knowledge within a Large Language Model is not stored neatly but is 'tangled up' with other cultures. The paper argues that biases are woven into the model's internal structure, meaning simple output grading is insufficient for understanding or fixing the problem.
- Culturescope
- This is a method used in the paper to peek inside the model's hidden layers. It acts like a 'tiny camera' inserted into the processing pipeline, allowing researchers to take snapshots of how cultural knowledge is stored and mixed within each layer of the model.
- Cultural Flattening Score (CF score)
- This metric quantifies how much one culture's unique knowledge is being overwritten or conflated with another culture's knowledge. It helps measure the degree to which a model generalizes or merges distinct cultural concepts into a single, less precise representation.
- Western-Dominance Bias
- The research found that when faced with ambiguous questions, the models tend to default to answers typical of high-resource, Western cultures. This bias is not random; it is deeply embedded in the model's internal processing and attention patterns.
Terminology used across episodes
This episode discusses
- Entangled in Representations: Mechanistic Investigation of Cultural Biases in Large Language Models · Paper Radio
- Challenges and Strategies in Cross-Cultural NLP
- CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages
The paper
Entangled in Representations: Mechanistic Investigation of Cultural Biases in Large Language Models · Read on arXiv
Haeun Yu, Arnav Arora, Seogyeong Jeong, Nadav Borenstein, Siddhesh Pawar, Jisu Shin, Jiho Jin, Junho Myung, Alice Oh, Isabelle Augenstein
University of Copenhagen · KAIST
The growing deployment of large language models (LLMs) across diverse cultural contexts necessitates a deeper understanding of LLMs' representations of different cultures. Prior work has focused on evaluating the cultural awareness of LLMs by only examining the text they generate. This approach overlooks the internal sources of cultural misrepresentation within the models themselves. To bridge this gap, we propose Culturescope, the first mechanistic interpretability-based method that probes the internal representations of different cultural knowledge in LLMs. We also introduce a cultural flattening score as a measure of the intrinsic cultural biases of the decoded knowledge from Culturescope. Additionally, we study how LLMs internalize cultural biases, which allows us to trace how cultural biases such as Western-dominance bias and cultural flattening emerge within LLMs. We find that low-resource cultures are less susceptible to cultural biases, likely due to the model's limited parametric knowledge. Our work provides a foundation for future research on mitigating cultural biases and enhancing LLMs' cultural understanding.
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Entangled in Representations: Mechanistic Investigation of Cultural Biases in Large Language Models".
Jane: The paper was written by Haeun Yu, Arnav Arora, Seogyeong Jeong, Nadav Borenstein, Siddhesh Pawar et al. from University of Copenhagen and KAIST.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back to the show, everyone. Today we're digging into a paper with a title that really sets the stage: "Entangled in Representations: Mechanistic Investigation of Cultural Biases in Large Language Models." Jane, when you first saw that word "entangled," what went through your mind?
Jane: Oh, Tom, it's perfect, honestly. It suggests that cultural biases aren't just something the model says, but something woven into its very fabric. The authors, from the University of Copenhagen and KAIST, are basically saying that if we want to fix cultural bias, we can't just look at the output text. We have to look inside the model's brain, so to speak.
Tom: Exactly. And that's a huge shift. Most prior work just asked the model questions and graded its answers. That's what they call extrinsic evaluation. But this paper is all about intrinsic evaluation—peeking at the hidden layers to see how cultural knowledge is actually stored and mixed together.
Jane: Right. And they introduce a method called Culturescope to do exactly that. It's like a tiny camera they insert into the model's processing pipeline to take snapshots of what's happening at each layer.
Lu: If I can jump in here, Tom. What excites me as a researcher is that they're not just saying "the model is biased." They're showing *where* that bias lives and *how* it flows. The title's "entangled" really captures the idea that knowledge about, say, Greece isn't stored in a neat little box. It's tangled up with knowledge about Turkey, about the UK, about everything else.
Meng: And from an engineering standpoint, that's both fascinating and terrifying. Because if the knowledge is entangled, then you can't just delete a bad memory or patch a single neuron. You have to untangle the whole mess.
Jane: That's a great point, Meng. And the paper actually gives us a way to measure that mess. They call it the Cultural Flattening score. It's a way to quantify how much one culture's distinctive knowledge is being overwritten or conflated with another culture's.
Tom: So it's not just "the model is biased," it's "the model is flattening Greek culture into Turkish culture" or something like that. That's a much more precise diagnosis.
Lu: Precisely. And that precision is what we need if we're ever going to build models that genuinely respect cultural diversity. We need to know the specific failure modes, not just a general sense of unfairness.
Meng: I'm curious about the practical side. How do they even get the model to "show" them its internal knowledge? That sounds like a technical nightmare.
Jane: That's the clever part, Meng. They use a technique called activation patching. They basically take the hidden state from one part of the model and inject it into another part, using a special prompt that asks the model to list cultural concepts. It's like asking the model to describe a picture it's holding in its mind, even if it can't show you the picture directly.
Tom: And the results are pretty striking. They found that the model's internal knowledge space is asymmetric. For example, knowledge about China flows into South Korea, but not the other way around. That's cultural flattening in action.
Lu: It's a powerful visualization. You can literally see the direction of the bias. And it's not random. It follows geographic and cultural proximity, which makes sense given the training data.
Meng: So the model is basically learning that "nearby cultures are interchangeable." That's a pretty deep-seated bias to have.
Jane: It is. And that's why this paper is so important. It's not just a new evaluation metric. It's a new way of seeing the problem. And once you can see it, you can start to fix it.
Tom: And that's exactly where we're headed next. We're going to talk about what they actually found when they put these models to the test.
Summary: Tom: So, Jane, we've set the stage with the title and the core idea. Now let's get into the meat of the paper. What did they actually do, and what did they find?
Jane: Well, Tom, they ran their Culturescope method on three different open-source models—Llama-three point one, aya-expanse, and Qwen2 point 5. And they used two cultural datasets, BLEnD and CAMeL-two covering a bunch of countries.
Lu: And the key finding, which I think is really profound, is that the models show a clear Western-dominance bias. When they're unsure, they tend to fall back on answers that are typical of high-resource, Western cultures.
Meng: So it's not just that they get things wrong. They get things wrong in a very specific, biased way. They're not making random mistakes.
Jane: Exactly. And they designed a clever experiment to prove it. They created multiple-choice questions with "hard negatives." So instead of giving the model one right answer and three obviously wrong ones, they gave it three plausible answers from different cultures.
Tom: Right. So for a question about China, they might offer the correct Chinese answer, but also the correct Korean answer and the correct American answer. That forces the model to actually distinguish between cultures, not just pick the most common answer.
Lu: And when the models got those questions wrong, they didn't pick randomly. They consistently picked the answer from the high-resource culture. That's the Western-dominance bias showing up in the output.
Meng: But here's the twist. They also looked at the attention maps—the internal focus of the model—and found the same pattern. The model's attention was literally drawn to the Western option, even before it made a mistake.
Jane: That's the really cool part. It means the bias isn't just a surface-level artifact. It's deeply embedded in how the model processes information. The model is internally "leaning" toward the Western answer.
Lu: And that's what they mean by "entangled in representations." The bias is in the weights, in the activations, in the attention patterns. It's everywhere.
Meng: So what about the cultural flattening score? Did that show the same thing?
Jane: It showed something related but distinct. The CF score measures how much one culture's distinctive knowledge is being used to represent another culture. And they found asymmetric connections. For example, China's knowledge flows into South Korea, but not the reverse.
Lu: And that's not just a random pattern. It's correlated with resource levels and geographic proximity. The model is essentially saying, "These cultures are similar, so I'll just use one to stand in for the other."
Tom: But here's a fascinating wrinkle they found. For low-resource cultures, like Assam or Azerbaijan, the model showed *less* bias. It didn't flatten them as much.
Meng: Wait, that sounds like good news. The model is less biased against them?
Jane: Not exactly. The paper argues it's because the model has so little knowledge about those cultures in the first place. It's not that it's being fair. It's that it has nothing to flatten *with*. There's no knowledge to be biased about.
Lu: That's a crucial distinction. The absence of bias isn't the same as the presence of fairness. It's just an absence of knowledge. And that's a very different problem to solve.
Tom: So we have two problems: one is cultural flattening, where the model overgeneralizes, and the other is cultural erasure, where the model just doesn't know anything. And they're both bad.
Meng: And they probably need different solutions. For flattening, you need to debias the model. For erasure, you need to add more data.
Jane: Exactly. And that's a big takeaway from this paper. It's not a one-size-fits-all problem. We need to understand the specific failure mode before we can fix it.
Tom: And that's what we're going to dig into next—what the paper suggests we do about all this.
Improvements: Tom: Alright, so we've established that the models are biased in specific ways. But what does this paper actually propose we *do* about it? Jane, what's the path forward?
Jane: Well, Tom, the paper doesn't give us a magic fix. But it does give us a roadmap. The first step is diagnosis. And that's what Culturescope is for. It's a tool to identify exactly where and how cultural bias is embedded in a model.
Lu: And that's more valuable than it sounds. Right now, if you want to make a model more culturally aware, you're kind of flying blind. You might add more data, or fine-tune on a specific culture, but you don't know if it's actually working on the right thing.
Meng: Right. It's like trying to fix a car engine without being able to open the hood. Culturescope opens the hood and shows you exactly which part is misfiring.
Jane: And once you can see that, you can start to think about targeted interventions. If you know that the model is flattening China into South Korea, you can specifically work on separating those two knowledge spaces.
Lu: The paper also suggests that we need different strategies for different resource levels. For high-resource cultures, the problem is overgeneralization. You need to debias, to untangle the knowledge. For low-resource cultures, the problem is underrepresentation. You need to add knowledge, not just debias.
Tom: So it's not just "make the model more diverse." It's "make the model more diverse in the right way, for the right reasons."
Meng: And that's a much more nuanced and practical approach. It acknowledges that cultural bias isn't a single problem. It's a collection of related problems that need different solutions.
Lu: Exactly. And the paper's cultural flattening score is a step toward making that distinction measurable. You can track whether your intervention is actually reducing flattening, or just making the model more confused.
Jane: Another interesting implication is for evaluation. This paper shows that just looking at accuracy on cultural questions isn't enough. You need to look at *how* the model is wrong, not just *that* it's wrong.
Tom: That's a big deal for the whole field of cultural AI evaluation. We need to move beyond simple benchmarks and start looking at the internal mechanisms.
Meng: I'm also thinking about the practical engineering side. These interpretability techniques are powerful, but they're also computationally expensive. Is this something that can be used in a real-world development pipeline?
Jane: That's a fair question. The paper uses eight-billion parameter models, which are manageable. But scaling this to larger models, or to many more cultures, would be a challenge.
Lu: But I think the value is more in the research direction. It tells us *what* to look for. Once we know that attention patterns are a reliable indicator of bias, we can develop cheaper proxies for that.
Tom: So it's not necessarily a tool you'd run on every model in production. But it's a tool you'd use to understand your model, to validate your training data, and to guide your fine-tuning.
Meng: That makes sense. It's a diagnostic tool, not a runtime tool. Like an MRI for the model.
Jane: Exactly. And the paper also validates Culturescope itself by showing that the cultural knowledge it extracts actually helps the model answer questions better. So it's not just an interesting toy. It's genuinely capturing useful information.
Lu: That's a crucial validation. It means the knowledge they're extracting from the internal layers is real and relevant, not just noise.
Tom: So the path forward is clear: use tools like Culturescope to diagnose, use the CF score to measure, and then develop targeted interventions based on the specific failure mode. That's a much more sophisticated approach than we've had before.
Jane: And it's a foundation for future work. The paper is very clear that this is just the beginning. There's a lot more to explore.
Meng: I'm curious to see how this scales to other languages and cultures, and whether the findings hold up in more diverse settings.
Tom: And that's a perfect segue into our final segment, where we wrap up and look at the big picture.
Conclusion: Tom: So, we've covered a lot of ground on "Entangled in Representations: Mechanistic Investigation of Cultural Biases in Large Language Models." Let's bring it all together.
Jane: The core message, Tom, is that cultural bias in LLMs is not a surface problem. It's deeply embedded in the model's internal representations. And to understand it, we need to look inside, not just at the outputs.
Lu: And the paper gives us the tools to do that. Culturescope lets us probe the internal knowledge, and the Cultural Flattening score lets us measure how cultures are being conflated.
Meng: The findings are clear: models show a Western-dominance bias, and they flatten low-resource cultures into their neighbors. But the paper also shows that low-resource cultures are less affected, simply because the model knows less about them.
Jane: That's the key nuance. It's not that the model is fair to low-resource cultures. It's that it's ignorant of them. And that's a different problem that needs a different solution.
Tom: So the takeaway is that we need targeted interventions. For high-resource cultures, we need to untangle the knowledge. For low-resource cultures, we need to add knowledge.
Lu: And we need to evaluate our models not just on accuracy, but on the *patterns* of their mistakes. This paper provides a framework for doing exactly that.
Meng: From an engineering perspective, this is a valuable diagnostic tool. It's not something you'd run in production, but it's something you'd use during development to understand your model's cultural blind spots.
Jane: And that's the exciting part. This isn't just an academic exercise. It has real implications for building AI systems that are more respectful of cultural diversity.
Tom: Absolutely. And with that, we're going to wrap up our discussion of this paper. It's a thought-provoking piece of research that opens up a whole new avenue for understanding and mitigating cultural bias in AI.
Jane: Thanks for joining us, everyone. We'll be back soon with another paper to dig into.
Tom: Until then, keep questioning, keep exploring, and keep looking inside the black box. Goodbye, everyone.
More episodes
- 2610.10857-Self-Supervised Keyframe Discovery for Horizon-Invariant Behavior Cloning
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization