AdaCultureSafe: Adaptive Cultural Safety Grounded by Cultural Knowledge in Large Language Models
cs.CL, cs.AI
Submitted: 2026-03-09
Updated: 2026-09-08
Comments: Accepted by EMNLP 2026 Findings
Project page: https://cs.mfa.gov.cn/zggmcg/ljmdd
License: http://creativecommons.org/licenses/by/4.0/
The gist: With the global proliferation of Large Language Models (LLMs), cultural safety, defined as the ability to generate respectful and appropriate responses across diverse cultures, becomes critical for
Terminology
Abstract
With the global proliferation of Large Language Models (LLMs), cultural safety, defined as the ability to generate respectful and appropriate responses across diverse cultures, becomes critical for responsible AI deployment. However, existing research often treats cultural safety and cultural knowledge in isolation. It remains unclear whether cultural safety is grounded in understanding varying cultural knowledge to enable LLMs to adaptively yield respectful and appropriate responses across diverse cultures, which trigger ethical concerns in cross-cultural scenarios. In this work, we introduce AdaCultureSafe, a dataset designed to jointly evaluate cultural safety and knowledge. Through AdaCultureSafe, we reveal a critical insight: a significant decoupling exists between cultural safety and cultural knowledge proficiency. Although LLMs possess rich cultural knowledge, they fail to leverage it to improve cultural safety. LLMs tend to rely on generic safety rather than safety based on culture-specific knowledge. Motivated by this, we propose a knowledge-grounded method that elicits internal cultural knowledge in LLMs during response generation. Experimental results demonstrate that our approach effectively improves cultural safety. Our work can aid better understanding the landscape of cultural safety of LLMs.
Sources
- The Cultural Gene of Large Language Models: A Study on the Impact of Cross-Corpus Training on Model Values and Biases
- Massively Multi-Cultural Knowledge Acquisition & LM Benchmarking
- Mistral 7B
- Layerwise Recall and the Geometry of Interwoven Knowledge in LLMs
- Multimodal Cultural Safety: Evaluation Framework and Alignment Strategies
- Qwen2.5 Technical Report
- Demystifying the Roles of LLM Layers in Retrieval, Knowledge, and Reasoning
- An Evaluation of Cultural Value Alignment in LLM
- AnswerCarefully: A Dataset for Improving the Safety of Japanese LLM Output
- Unraveling LLM Jailbreaks Through Safety Knowledge Neurons
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering