Robustness of Vision Language Models Against Split-Image Harmful Input Attacks
cs.CV, cs.AI
Submitted: 2026-02-08
Updated: 2026-08-30
Comments: Accepted in ACM CCS 2026. 26 Pages, long conference paper
License: http://creativecommons.org/licenses/by/4.0/
The gist: Vision-Language Models (VLMs) are now a core part of modern AI.
Terminology
Abstract
Vision-Language Models (VLMs) are now a core part of modern AI. Recent work proposed several visual jailbreak attacks using single/ holistic images. However, contemporary VLMs demonstrate strong robustness against such attacks due to extensive safety alignment through preference optimization, e.g., reinforcement learning from human feedback (RLHF). In this work, we identify a new vulnerability: while VLM pretraining and instruction tuning generalize well to split-image inputs, safety alignment is typically performed only on holistic images and does not account for harmful semantics distributed across multiple image fragments. Consequently, VLMs often fail to detect and reject harmful split-image inputs, in which unsafe cues emerge only upon combining images. We introduce novel split-image visual jailbreak attacks (SIVA) that exploit this misalignment. Unlike prior optimization-based attacks, which exhibit poor black-box transferability due to architectural and prior mismatches across models, our attacks evolve in progressive phases from naive splitting to an adaptive white-box attack, culminating in a black-box transfer attack. Our strongest strategy leverages a novel adversarial knowledge distillation(Adv-KD) algorithm to substantially improve cross-model transferability. Evaluations on four state-of-the-art modern VLMs and three jailbreak datasets demonstrate that our strongest attack achieves up to 44% higher transfer success than existing baselines. Lastly, we propose efficient ways to address this critical vulnerability in the current VLM safety alignment.
Sources
- Pixtral 12B
- Detecting Language Model Attacks with Perplexity
- Qwen3-VL Technical Report
- ETA: Evaluating Then Aligning Safety of Vision Language Models at Inference Time
- E$^2$AT: Multimodal Jailbreak Defense via Dynamic Joint Optimization for Multimodal Large Language Models
- JailBreakV: A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak Attacks
- Towards Deep Learning Models Resistant to Adversarial Attacks
- Chain-of-Thought Driven Adversarial Scenario Extrapolation for Robust Language Models
- Adversarial Manipulation of Deep Representations
- Proximal Policy Optimization Algorithms
- LLaVA-MoD: Making LLaVA Tiny via MoE Knowledge Distillation
- LLAVADI: What Matters For Multimodal Large Language Models Distillation
- MM-RLHF: The Next Step Forward in Multimodal LLM Alignment
- BlueSuffix: Reinforced Blue Teaming for Vision-Language Models Against Jailbreak Attacks
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
- Safety Fine-Tuning at (Almost) No Cost: A Baseline for Vision Large Language Models
- Universal and Transferable Adversarial Attacks on Aligned Language Models
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models