MMJailBench: A Factorized Benchmark for Disentangling Multimodal Jailbreak Vulnerabilities
cs.CR, cs.AI, cs.MM
Submitted: 2026-08-26
Updated: 2026-08-26
License: http://creativecommons.org/licenses/by/4.0/
The gist: Multimodal Large Language Models (MLLMs) are increasingly deployed in real-world applications, yet how different factors shape their jailbreak vulnerabilities remains poorly understood.
Terminology
Abstract
Multimodal Large Language Models (MLLMs) are increasingly deployed in real-world applications, yet how different factors shape their jailbreak vulnerabilities remains poorly understood. Existing benchmarks often couple harmful intent, prompt framing, visual semantics, and instruction carrier within individual jailbreak instances, obscuring the specific sources of observed vulnerabilities. To address this limitation, we introduce MMJailBench, a factorized benchmark that systematically varies and combines these factors under controlled configurations, enabling fine-grained comparison and factor-level attribution. Large-scale evaluations across 16 open-weight and proprietary MLLMs reveal highly heterogeneous and model-dependent vulnerability profiles. Jailbreak vulnerability varies markedly across harm domains, exposing uneven coverage in current multimodal safety alignment. Prompt framing emerges as the dominant source of variation, task-relevant visual semantics systematically increase jailbreak susceptibility with authority-like cues exposing particularly pronounced vulnerabilities, and visually rendered instructions do not consistently increase jailbreak susceptibility relative to direct textual instructions. To further investigate the risks introduced by multimodal context, we conduct diagnostic analyses on a representative open-weight model and identify vulnerability-associated patterns in internal representations and cross-modal interactions. Finally, we develop a modular multimodal jailbreak evaluation suite with full and lightweight configurations, multiple judge options, and multidimensional metrics, enabling reproducible, scalable, and cost-efficient multimodal jailbreak auditing.
Sources
- GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
- Baseline Defenses for Adversarial Attacks Against Aligned Language Models
- LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
- OmniSafeBench-MM: A Unified Benchmark and Toolbox for Multimodal Jailbreak Attack-Defense Evaluation
- Qwen3-VL Technical Report
- Qwen2.5-VL Technical Report
- JailBreakV: A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak Attacks
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
- Omni-SafetyBench: A Benchmark for Safety Evaluation of Audio-Visual Large Language Models
- Universal and Transferable Adversarial Attacks on Aligned Language Models
- OpenAI GPT-5 System Card
- Gemini: A Family of Highly Capable Multimodal Models
- Gemma 3 Technical Report
- Kimi K2.5: Visual Agentic Intelligence
- Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs
- STEP3-VL-10B Technical Report
- Ministral 3
- A StrongREJECT for Empty Jailbreaks
- Qwen3 Technical Report
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs