SGM: Safety Glasses for Multimodal Large Language Models via Neuron-Level Detoxification
cs.CL, cs.AI
Submitted: 2025-12-17
Updated: 2026-08-31
Comments: Accepted to the EMNLP 2026 Main Conference
License: http://creativecommons.org/licenses/by/4.0/
The gist: Disclaimer: Samples in this paper may be harmful and cause discomfort.
Terminology
Abstract
Disclaimer: Samples in this paper may be harmful and cause discomfort. Multimodal large language models (MLLMs) enable multimodal understanding but inherit toxic signals from weakly curated pretraining corpora, leading to explicitly toxic outputs, especially under adversarial triggers that late, opaque training-free detoxification methods struggle to handle. We propose SGM, a white-box neuron-level multimodal intervention that acts like safety glasses for toxic neurons: it recalibrates a set of toxicity-associated neurons via expertise-weighted soft suppression, neutralizing harmful cross-modal activations without any parameter updates. We establish MM-TOXIC-QA, a multimodal toxicity data framework, and compare SGM with existing detoxification techniques. Experiments on open-source MLLMs show that SGM mitigates explicit toxicity in standard and adversarial conditions, cutting average harmful rates from 45.0% to 4.5% while preserving fluency and multimodal reasoning. SGM is extensible, and its combined defenses, denoted as SGM*, integrate with existing detoxification methods for stronger safety performance, providing an interpretable, low-cost solution for toxicity-controlled multimodal generation.
Sources
- Fairness and Bias in Multimodal AI: A Survey
- Unraveling and Mitigating Safety Alignment Degradation of Vision-Language Models
- Gaining Wisdom from Setbacks: Aligning Large Language Models via Mistake Analysis
- MirrorCheck: Efficient Adversarial Defense for Vision-Language Models
- CoCA: Regaining Safety-awareness of Multimodal Large Language Models with Constitutional Calibration
- Safe RLHF-V: Safe Reinforcement Learning from Multi-modal Human Feedback
- GPT-4 Technical Report
- MLLM-Protector: Ensuring MLLM's Safety without Hurting Performance
- SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks
- Whispering Experts: Neural Interventions for Toxicity Mitigation in Language Models
- Self-conditioning pre-trained language models
- Safety Assessment of Chinese Large Language Models
- Safeguarding Vision-Language Models Against Patched Visual Prompt Injectors
- PclGPT: A Large Language Model for Patronizing and Condescending Language Detection
- Cross-modality Information Check for Detecting Jailbreaking in Multimodal Large Language Models
- SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware Decoding
- BlueSuffix: Reinforced Blue Teaming for Vision-Language Models Against Jailbreak Attacks
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering