Transfer Safety Awareness for Cross-Modal Safety Drift in Multimodal Large Language Models
cs.MM, cs.AI, cs.CL, cs.CR
Submitted: 2026-09-02
Updated: 2026-09-03
Comments: Accepted to Findings of EMNLP 2026
Code: https://github.com/cucu220123/safety-awareness
License: http://creativecommons.org/licenses/by/4.0/
The gist: Visual modality enhances the capabilities of multimodal large language models (MLLMs) but also introduces a safety concern: a benign textual query may convey harmful intent when grounded in a visual
Terminology
Abstract
Visual modality enhances the capabilities of multimodal large language models (MLLMs) but also introduces a safety concern: a benign textual query may convey harmful intent when grounded in a visual image. We term this cross-modal safety drift and our pilot studies show that the safety response rate for such requests is substantially lower than that for requests containing explicitly unsafe text. This paper aims to systematically study this issue. First, we conduct an empirical analysis to identify representative unsafe response patterns. Building on these, we interpret model representations and attentions, revealing that visually risky cues receive limited attention and weakly trigger refusal. Motivated by the observation that safety signals from unsafe text processing can be transferred, we propose safety-awareness representation transfer (SRT), a lightweight direction-refinement method that mitigates cross-modal safety drift with a frozen MLLM backbone. Experiments across multiple benchmarks and models show that SRT effectively improves safety in diverse cross-modal settings while preserving utility. Code is available at https://github.com/cucu220123/safety-awareness.
Sources
- Attention Misses Visual Risk: Risk-Adaptive Steering for Multimodal Safety Alignment
- Steering Language Models With Activation Engineering
- Risk Awareness Injection: Calibrating Vision-Language Models for Safety without Compromising Utility
Related papers
- ControlFoley: Unified and Controllable Video-to-Audio Generation with Cross-Modal Conflict Handling
- Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models
- Mitigating GenAI-powered Evidence Pollution for Out-of-Context Multimodal Misinformation Detection
- A Rate-Distortion-Classification Approach for Lossy Image Compression