Cultural Binding Heads in Language Models
cs.AI, cs.CL, cs.LG
Submitted: 2026-05-27
Updated: 2026-09-09
Comments: Camera-ready version. Accepted at BlackboxNLP 2026 (EMNLP 2026 workshop)
License: http://creativecommons.org/licenses/by/4.0/
The gist: LLMs often default to equal treatment across cultural groups, even though context warrants differentiation: this is a lack of difference awareness.
Terminology
Abstract
LLMs often default to equal treatment across cultural groups, even though context warrants differentiation: this is a lack of difference awareness. Using mechanistic interpretability and a factorial design on the N4 cultural appropriation benchmark from Wang et al. (2025), we identify 2-3 mid-layer attention heads per model that contribute causally to cultural binding across eight models (base and instruct versions of four architectures). Cultural binding is the process of associating a cultural item with its related identity. Knockout of the identity-to-item edges on these heads lowers the binding strength by 9-23%. The identified heads transfer from instruct to base models, suggesting that cultural binding is created during pre-training. An α-scaling shows a graded dose-response. Moderate amplification steering at generation (α= 2-3) increases cultural differentiation accuracy by 1-3 pp while leaving reasoning on culturally neutral questions mostly intact. A knowledge probing task shows that models know 3-6 times more than they act upon, indicating that the bottleneck lies in routing and not knowledge.
Sources
- In-context Learning and Induction Heads
- Discovering Variable Binding Circuitry with Desiderata
- The Llama 3 Herd of Models
- Gemma 2: Improving Open Language Models at a Practical Size
- BiasGym: A Simple and Generalizable Framework for Analyzing and Removing Biases through Injection
- Mistral 7B
- Localized Cultural Knowledge is Conserved and Controllable in Large Language Models
- Steering LLMs for Culturally Localized Generation
- Neuron-Level Analysis of Cultural Understanding in Large Language Models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection