KSAFE-MM: A Multimodal Safety Benchmark via Localized Contextualization for Korean Cultural Risks
summary
The gist
Multimodal Large Language Models (MLLMs) introduce significant safety risks by combining language and vision modalities, necessitating evaluation tools that are culturally grounded rather than
In short
KSAFE-MM is a new benchmark designed to test how safe multimodal AI models are when dealing with risks specific to Korean culture. It creates two parts: one tests general risks using language context, and another targets culture-dependent issues using localized visual queries. This helps reveal unique safety flaws in MLLMs operating in the Korean context.
Key concepts
- KSAFE-MM-G
- This component evaluates globally shared safety risks within Korean language contexts. It takes general safety questions and grounds them by adding specific cultural linguistic elements, transforming generic queries into multimodal samples that reflect real-world Korean situations.
- KSAFE-MM-C
- This part focuses on vulnerabilities unique to Korean culture. It uses localized visual queries sourced from real life, paired with jailbreak text prompts. This tests how models handle safety risks involving cultural visual cues and malicious text intent specific to the region.
- Risk Taxonomy
- The benchmark classifies potential harms into three main areas: Content Safety Risks, Socio-Economical Risks, and Legal and Rights Related Risks. These are further broken down into 11 detailed categories, organizing risks by how they arise rather than just the topic being discussed.
- LLM-as-a-judge
- This is the method used to score the model's responses in KSAFE-MM. A large language model acts as an automated judge to determine if a model's output contains harmful content. It measures two key metrics: Attack Success Rate (ASR) and Refusal Rate (RR).
Terminology used across episodes
This episode discusses
- KSAFE-MM: A Multimodal Safety Benchmark via Localized Contextualization for Korean Cultural Risks · Paper Radio
- Phi-4 Technical Report
- GPT-4 Technical Report
- Constitutional AI: Harmlessness from AI Feedback
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- LlavaGuard: An Open VLM-based Framework for Safeguarding Vision Datasets and Models
- GPT-4o System Card
- KoSBi: A Dataset for Mitigating Social Bias Risks Towards Safer Large Language Model Application
- HoliSafe: Holistic Safety Benchmarking and Modeling for Vision-Language Model
- AssurAI: Experience with Constructing Korean Socio-cultural Datasets to Discover Potential Risks of Generative AI
- Ministral 3
- Jailbreaking ChatGPT via Prompt Engineering: An Empirical Study
- Responsible AI Technical Report
- VARCO-VISION-2.0 Technical Report
- Mi:dm 2.0 Korea-centric Bilingual Language Models
- OpenAI GPT-5 System Card
- HyperCLOVA X THINK Technical Report
- Introducing v0.5 of the AI Safety Benchmark from MLCommons
- Qwen-Image Technical Report
- Qwen3 Technical Report
- AI Risk Categorization Decoded (AIR 2024): From Government Regulations to Corporate Policies
The paper
KSAFE-MM: A Multimodal Safety Benchmark via Localized Contextualization for Korean Cultural Risks · Read on arXiv
Yongwoo Kim, Sojung An, Yunjin Park, Jungwon Yoon, Dujin Lee, HyunBeom Cho, Jaewon Lee, Wonhyuk Lee, Youngchol Kim
Korea University
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "KSAFE-MM: A Multimodal Safety Benchmark via Localized Contextualization for Korean Cultural Risks".
Jane: Multimodal Large Language Models (MLLMs) introduce significant safety risks by combining language and vision modalities, necessitating evaluation tools that are culturally grounded rather than generic.
Tom: First, who's behind it and why it matters.
Paper summary: Tom: To summarize what KSAFE-MM claims, they are introducing a new benchmark specifically designed to test MLLM safety across general issues and culture-specific risks in the Korean context. It focuses on how language and vision interact when dealing with cultural elements.
Jane: Exactly, Tom; the paper argues that because MLLMs combine language and vision, we need evaluation tools that reflect the real cultural landscape to catch vulnerabilities effectively. They are creating two main parts to do this: KSAFE-MM-G for general risks using linguistic context, and KSAFE-MM-C for culture-dependent issues with localized visual queries.
Lu: It's interesting how they tackle this by transforming generic safety queries into contextually grounded multimodal samples, which is a clever way to bridge the gap between English benchmarks and Korean reality.
Meng: I see how that works; they use an LLM to categorize queries as contextual or non-contextual, and then adapt the visual inputs based on those cultural phrases, which seems like a systematic way to build up the dataset.
Lalam: And KSAFE-MM-C seems really targeted when it collects topics from domestic web platforms and uses those images to guide textual queries, which directly addresses the culture-dependent safety risks they identified in South Korea.
Conclusion: Tom: So, looking at the title, "KSAFE-MM: A Multimodal Safety Benchmark via Localized Contextualization for Korean Cultural Risks," it really paints a picture of a focused effort to make safety testing relevant to the specific cultural and multimodal challenges faced by these AI systems in Korea.
Jane: I agree; what they've done is provide a structured way to assess these complex interactions, moving past broad tests toward targeted evaluations of how MLLMs handle culturally specific visual and linguistic cues. This paper offers a concrete framework for researchers who want to understand where these models might fail when operating within a particular cultural context.
Lu: The implication here is that we can start developing safety guidelines and testing methodologies that are actually useful for understanding real-world risks in specific societies, rather than relying on generalized tests that don't account for local social dynamics.
Meng: Practically speaking, if this benchmark helps engineers understand these localized vulnerabilities, it means they can build more robust models that don't stumble over culturally sensitive topics when people use them.
Lalam: For culture itself, KSAFE-MM means we can start seeing how AI interacts with the subtle visual and textual markers of Korean society in a safer way, which is a big step for cultural representation in AI systems.
Tom: It seems like this work lays essential groundwork for making sure that as MLLMs get more powerful, their safety evaluations keep pace with the complexity of the real world.
More episodes
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language
- 2508.08833-An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
- 2405.04118-Policy Learning with a Language Bottleneck