Probabilistic Modeling of Jailbreak on Multimodal LLMs: From Quantification to Application
cs.CR, cs.CV
Submitted: 2025-03-10
Updated: 2026-09-22
Comments: ESORICS 2026
License: http://creativecommons.org/licenses/by/4.0/
The gist: Recently, Multimodal Large Language Models (MLLMs) have demonstrated their superior ability in understanding multimodal content.
Terminology
Abstract
Recently, Multimodal Large Language Models (MLLMs) have demonstrated their superior ability in understanding multimodal content. However, they remain vulnerable to jailbreak attacks, which exploit weaknesses in their safety alignment to generate harmful responses. Previous studies categorize jailbreaks as successful or failed based on whether responses contain malicious content. However, given the stochastic nature of MLLM responses, this binary classification of an input's ability to jailbreak MLLMs is inappropriate. Derived from this viewpoint, we introduce jailbreak probability to quantify the jailbreak potential of an input, which represents the likelihood that MLLMs generated a malicious response when prompted with this input. We approximate this probability through multiple queries to MLLMs. After modeling the relationship between input hidden states and their corresponding jailbreak probability using Jailbreak Probability Prediction Network (JPPN), we use continuous jailbreak probability for optimization. Specifically, we propose Jailbreak-Probability-based Attack (JPA) that optimizes adversarial perturbations on input image to maximize jailbreak probability, and further enhance it as Multimodal JPA (MJPA) by including monotonic text rephrasing. To counteract attacks, we also propose Jailbreak-Probability-based Finetuning (JPF), which minimizes jailbreak probability through MLLM parameter updates. Extensive experiments show that (1) (M)JPA yields significant improvements when attacking a wide range of models under both white and black box settings. (2) JPF vastly reduces jailbreaks by at most over 60%. Both of the above results demonstrate the significance of introducing jailbreak probability to make nuanced distinctions among input jailbreak abilities.
Sources
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
- Safe RLHF: Safe Reinforcement Learning from Human Feedback
- How Robust is Google's Bard to Adversarial Image Attacks?
- FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts
- Explaining and Harnessing Adversarial Examples
- LlavaGuard: An Open VLM-based Framework for Safeguarding Vision Datasets and Models
- A Survey of Safety and Trustworthiness of Large Language Models through the Lens of Verification and Validation
- GPT-4o System Card
- Adam: A Method for Stochastic Optimization
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
- DeepInception: Hypnotize Large Language Model to Be Jailbreaker
- Images are Achilles' Heel of Alignment: Exploiting Visual Vulnerabilities for Jailbreaking Multimodal Large Language Models
- Safety of Multimodal Large Language Models on Images and Texts
- DeepSeek-VL: Towards Real-World Vision-Language Understanding
- Groma: Localized Visual Tokenization for Grounding Multimodal Large Language Models
- A StrongREJECT for Empty Jailbreaks
- Jailbreak Attacks and Defenses Against Large Language Models: A Survey
- A Survey on Multimodal Large Language Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs