MMAligner: Safeguarding Multimodal Large Language Models through Representation Calibration
Shenyi Zhang, Keyan Guo, Zihao Wang, Xuebin Li, Lingchen Zhao, Hongxin Hu, Chao Shen, Qian Wang
cs.CR
Submitted: 2026-08-06
Comments: To Appear in the Proceedings of The ACM Conference on Computer and Communications Security (CCS), 2026
Code: https://github.com/shenyizg/MMAligner
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- The Llama 3 Herd of Models
- HiddenDetect: Detecting Jailbreak Attacks against Large Vision-Language Models via Monitoring Hidden States
- Unraveling and Mitigating Safety Alignment Degradation of Vision-Language Models
- Safety Alignment for Vision Language Models
- GPT-4 Technical Report
- Llama Guard 3 Vision: Safeguarding Human-AI Image Understanding Conversations
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- VLMGuard: Bootstrapping Malicious Prompt Detectors from Unlabeled Vision-Language Prompts in the Wild
- A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations
- A Guardrail for Safety Preservation: When Safety-Sensitive Subspace Meets Harmful-Resistant Null-Space
- Universal and Transferable Adversarial Attacks on Aligned Language Models
- Image-to-Text Logic Jailbreak: Your Imagination can Help You Do Anything
- VITA-Audio: Fast Interleaved Cross-Modal Token Generation for Efficient Large Speech-Language Model
- When Helpers Become Hazards: A Benchmark for Analyzing Multimodal LLM-Powered Safety in Daily Life
- X-Boundary: Establishing Exact Safety Boundary to Shield LLMs from Multi-Turn Jailbreaks without Compromising Usability
- Jailbreaking Attack against Multimodal Large Language Model
- Heuristic-Induced Multimodal Risk Distribution Jailbreak Attack for Multimodal Large Language Models
- SoK: Evaluating Jailbreak Guardrails for Large Language Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs