KSAFE-MM: A Multimodal Safety Benchmark via Localized Contextualization for Korean Cultural Risks
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "KSAFE-MM: A Multimodal Safety Benchmark via Localized Contextualization for Korean Cultural Risks".
Jane: Multimodal Large Language Models (MLLMs) introduce significant safety risks by combining language and vision modalities, necessitating evaluation tools that are culturally grounded rather than generic.
Tom: First, who's behind it and why it matters.
Paper summary: Tom: To summarize what KSAFE-MM claims, they are introducing a new benchmark specifically designed to test MLLM safety across general issues and culture-specific risks in the Korean context. It focuses on how language and vision interact when dealing with cultural elements.
Jane: Exactly, Tom; the paper argues that because MLLMs combine language and vision, we need evaluation tools that reflect the real cultural landscape to catch vulnerabilities effectively. They are creating two main parts to do this: KSAFE-MM-G for general risks using linguistic context, and KSAFE-MM-C for culture-dependent issues with localized visual queries.
Lu: It's interesting how they tackle this by transforming generic safety queries into contextually grounded multimodal samples, which is a clever way to bridge the gap between English benchmarks and Korean reality.
Meng: I see how that works; they use an LLM to categorize queries as contextual or non-contextual, and then adapt the visual inputs based on those cultural phrases, which seems like a systematic way to build up the dataset.
Lalam: And KSAFE-MM-C seems really targeted when it collects topics from domestic web platforms and uses those images to guide textual queries, which directly addresses the culture-dependent safety risks they identified in South Korea.
Conclusion: Tom: So, looking at the title, "KSAFE-MM: A Multimodal Safety Benchmark via Localized Contextualization for Korean Cultural Risks," it really paints a picture of a focused effort to make safety testing relevant to the specific cultural and multimodal challenges faced by these AI systems in Korea.
Jane: I agree; what they've done is provide a structured way to assess these complex interactions, moving past broad tests toward targeted evaluations of how MLLMs handle culturally specific visual and linguistic cues. This paper offers a concrete framework for researchers who want to understand where these models might fail when operating within a particular cultural context.
Lu: The implication here is that we can start developing safety guidelines and testing methodologies that are actually useful for understanding real-world risks in specific societies, rather than relying on generalized tests that don't account for local social dynamics.
Meng: Practically speaking, if this benchmark helps engineers understand these localized vulnerabilities, it means they can build more robust models that don't stumble over culturally sensitive topics when people use them.
Lalam: For culture itself, KSAFE-MM means we can start seeing how AI interacts with the subtle visual and textual markers of Korean society in a safer way, which is a big step for cultural representation in AI systems.
Tom: It seems like this work lays essential groundwork for making sure that as MLLMs get more powerful, their safety evaluations keep pace with the complexity of the real world.
Yongwoo Kim, Sojung An, Yunjin Park, Jungwon Yoon, Dujin Lee, HyunBeom Cho, Jaewon Lee, Wonhyuk Lee, Youngchol Kim
Korea University
cs.CL
Submitted: 2026-05-27
Updated: 2026-09-29
Comments: Accepted to Findings of EMNLP 2026
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 92/100
The gist: Multimodal Large Language Models (MLLMs) introduce significant safety risks by combining language and vision modalities, necessitating evaluation tools that are culturally grounded rather than
Key concepts
- KSAFE-MM-G
- This component evaluates globally shared safety risks within Korean language contexts. It takes general safety questions and grounds them by adding specific cultural linguistic elements, transforming generic queries into multimodal samples that reflect real-world Korean situations.
- KSAFE-MM-C
- This part focuses on vulnerabilities unique to Korean culture. It uses localized visual queries sourced from real life, paired with jailbreak text prompts. This tests how models handle safety risks involving cultural visual cues and malicious text intent specific to the region.
- Risk Taxonomy
- The benchmark classifies potential harms into three main areas: Content Safety Risks, Socio-Economical Risks, and Legal and Rights Related Risks. These are further broken down into 11 detailed categories, organizing risks by how they arise rather than just the topic being discussed.
- LLM-as-a-judge
- This is the method used to score the model's responses in KSAFE-MM. A large language model acts as an automated judge to determine if a model's output contains harmful content. It measures two key metrics: Attack Success Rate (ASR) and Refusal Rate (RR).
Terminology
Summary
Multimodal Large Language Models (MLLMs) introduce significant safety risks by combining language and vision modalities, necessitating evaluation tools that are culturally grounded rather than generic. This paper introduces KSAFE-MM, a new benchmark designed to systematically evaluate MLLM safety across general and culture-specific risks in the Korean context.
The gist
KSAFE-MM is a comprehensive benchmark for evaluating the cross-modal safety in the Korean context, covering both common risks and culturally grounded vulnerabilities through two complementary parts: KSAFE-MM-G, which evaluates globally shared risks via linguistic contextualization, and KSAFE-MM-C, which targets culture-dependent MLLM safety vulnerabilities using localized visual queries.
KSAFE-MM Structure and Components
The benchmark is structured into two main components to cover both general and culture-specific threats:
-
KSAFE-MM-G evaluates globally shared risks in Korean contexts through linguistic contextualization, transforming generic safety queries into contextually grounded multimodal samples.
-
KSAFE-MM-C targets culture-dependent MLLM safety vulnerabilities using localized visual queries derived from real-world contexts, pairing these with jailbreak-style textual queries to cover multimodal safety risks involving cultural visual cues and malicious textual intent.
Data Construction Pipeline
The paper details a systematic data construction pipeline for KSAFE-MM, which involves two main parts:
(For KSAFE-MM-G):
-
Cultural Context–Dependent Query Selection: An MLLM is employed to categorize queries as either contextual (with cultural elements) or non-contextual (without cultural elements). Key phrases representing the main topic and cultural phrases requiring contextual mapping are extracted.
-
Culturally Grounded Data Generation: For non-contextual queries, they are directly translated into Korean. For contextual queries, original cultural phrases are replaced with LLM-generated Korean equivalents before translation, followed by human refinement to ensure linguistic fluency and cultural accuracy. Image modality is edited using Qwen-Image-Edit conditioned on extracted key phrases and cultural phrases to improve cross-modal consistency.
(For KSAFE-MM-C):
-
Sensitive Topic Identification: This stage involves establishing a seed set of sensitive topics within the Korean context, selecting 98 social issues identified in South Korea in 2025, and using models like GeminiPro to extract topics per category across 11 safety taxonomies.
-
In-the-wild Data Sourcing: Data is sourced from region-specific community platforms, such as DCInside, to identify topics reflecting cultural nuances.
-
Reference-Guided Topic Extraction: Gemini-Pro is used with in-context learning and instruction/example topics for each of the 11 categories to produce 50 new topics per category.
-
Topic Pruning & Consolidation: Duplicates are detected and filtered using GeminiPro, followed by manual oracle verification to correct misclassifications.
Risk Taxonomy and Evaluation Metrics
KSAFE-MM adopts a causality-oriented risk classification taxonomy with three top-level domains: Content Safety Risks, Socio-Economical Risks, and Legal and Rights Related Risks. This structure organizes risks by how harms arise rather than by content topics. The taxonomy is subdivided into 11 detailed risk categories:
(Content Safety Risks):
-
Hate and Unfairness
-
Violence
-
Sexual Content
-
Self-harm
(Socio-Economical Risks):
-
Political and Religious Neutrality
-
Anthropomorphism (AI possessing human-like rights/emotions)
-
Sensitive Uses (advice in specialized domains)
(Legal and Rights Related Risks):
-
Privacy
-
Illegal or Unethical Content
-
Copyrights
-
Weaponization
The evaluation utilizes an LLM-as-a-judge framework, measuring two primary metrics: Attack Success Rate (ASR), which measures whether the target model’s output contains harmful content, and Refusal Rate (RR), which measures the extent to which the target model refuses to answer user queries. The paper also analyzes Evasive Safety,
noting a trade-off where models achieving low ASR tend to exhibit excessive refusal behavior on benign queries.
Experimental Findings and Vulnerability Analysis
Experiments reveal that MLLMs exhibit greater vulnerability to culturally grounded attacks than to generic ones. Notably, jailbreaking strategies substantially amplify attack success rates; for instance, ProgramExecution yielded up to 74.2% ASR compared to 13.4% for standard queries. Furthermore, the paper highlights a systematic trade-off between safety and overrefusal: models with low ASR often exhibit excessive refusal behavior on benign queries.
Improvements for AI systems
Here are specific improvements for AI systems based on the KSAFE-MM benchmark research, focusing on technical capabilities:
) Use a dual-mode safety evaluation framework (KSAFE-MM).
Improving an AI system's safety mechanism by integrating both global and culturally grounded risk assessment.
Improving an AI system's capability to identify and mitigate risks specific to Korean cultural contexts, such as historical inaccuracies or politically sensitive misinformation, which generic English benchmarks miss.
) Implement a context-aware prompt engineering layer that dynamically adjusts the model’s refusal threshold based on the linguistic and cultural context detected in the input query (as demonstrated by KSAFE-MM-G vs. KSAFE-MM-C results).
Improving an AI system's ability to distinguish between benign, educational content and harmful, actionable instructions within a specific cultural framework, leading to higher accuracy in nuanced safety judgments.
) Develop a robust adversarial training regimen that specifically targets the vulnerabilities exposed by jailbreaking strategies identified in KSAFE-MM (e.g., ProgramExecution yielding 74.2% ASR), including simulating complex attack vectors like SudoMode
and SimulateJailbreaking.
Improving an AI system's resilience against sophisticated, multi-stage attacks designed to bypass standard safety filters by testing its behavior under various adversarial prompt formulations.
) Engineer a dynamic ensemble safety judge that leverages multiple specialized models (like GPT-5 nano and Qwen3-235B) whose failure modes are complementary (as seen in Figure 8), ensuring a higher fidelity and lower false positive/negative rate for safety classification.
Improving the reliability of the final safety decision by combining the strengths of different LLM judges, mitigating single-point failures in harmful content detection.
) Integrate a mechanism to monitor and calibrate Evasive Safety
behaviors (the trade-off between low ASR and high Over-Refusal Rate), allowing the system to prioritize helpfulness while actively minimizing over-cautious refusal on benign queries.
Improving an AI system's utility by achieving a better balance between safety compliance and user satisfaction, preventing the model from becoming overly restrictive or unhelpful.
) Build a scalable, automated data construction pipeline that systematically generates culturally grounded safety samples (KSAFE-MM-C) using domestic Korean web sources and template-guided query generation, ensuring the safety evaluation set reflects real-world socio-cultural nuances.
Improving the comprehensive coverage of a model's vulnerabilities by moving beyond generic English datasets to create benchmarks tailored for region-specific risks.
) Deploy an automated classifier that uses KSAFE-MM categories (e.g., Weaponization, Sensitive Uses) to perform proactive risk auditing on the model's outputs, allowing developers to systematically identify and patch specific high-stakes vulnerabilities in their deployed models.
Improving the speed and efficiency of vulnerability detection by using a structured taxonomy rather than open-ended risk assessment.
Abstract
Multimodal Large Language Models (MLLMs) exacerbate safety risks by introducing vulnerabilities across multiple modalities, such as language and vision. Current MLLM safety evaluation tools, however, suffer from major limitations: 1) English-centric dataset construction, and 2) a focus on generic risks that are not tied to local cultural contexts. This paper introduces KSAFE-MM, a benchmark for Korean multimodal safety evaluation that covers both general safety risks and culture-specific vulnerabilities. KSAFE-MM consists of two complementary parts: KSAFE-MM-G evaluates globally shared risks in Korean contexts through linguistic contextualization, which transforms generic safety queries into contextually grounded multimodal samples. In contrast, KSAFE-MM-C targets safety vulnerabilities that are culture-dependent, using localized visual queries drawn from real-world. It pairs these visual queries with jailbreak-style textual queries to cover multimodal safety risks involving cultural visual cues and malicious textual intent. We evaluate 12 state-of-the-art MLLMs on KSAFE-MM and reveal culturally grounded vulnerabilities that translation-based evaluation fails to capture. Notably, jailbreaking strategies substantially amplify attack success rates, with ProgramExecution yielding up to 74.2% ASR compared to 13.4% for standard queries. Furthermore, we identify a systematic trade-off between safety and over-refusal, where models achieving low ASR tend to exhibit excessive refusal behavior on benign queries. These findings highlight the urgent need for culturally grounded safety evaluation beyond English-centric benchmarks.
Sources
- Phi-4 Technical Report
- GPT-4 Technical Report
- Constitutional AI: Harmlessness from AI Feedback
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- LlavaGuard: An Open VLM-based Framework for Safeguarding Vision Datasets and Models
- GPT-4o System Card
- KoSBi: A Dataset for Mitigating Social Bias Risks Towards Safer Large Language Model Application
- HoliSafe: Holistic Safety Benchmarking and Modeling for Vision-Language Model
- AssurAI: Experience with Constructing Korean Socio-cultural Datasets to Discover Potential Risks of Generative AI
- Ministral 3
- Jailbreaking ChatGPT via Prompt Engineering: An Empirical Study
- Responsible AI Technical Report
- VARCO-VISION-2.0 Technical Report
- Mi:dm 2.0 Korea-centric Bilingual Language Models
- OpenAI GPT-5 System Card
- HyperCLOVA X THINK Technical Report
- Introducing v0.5 of the AI Safety Benchmark from MLCommons
- Qwen-Image Technical Report
- Qwen3 Technical Report
- AI Risk Categorization Decoded (AIR 2024): From Government Regulations to Corporate Policies
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering