Why Does Weak-OOD Help? A Further Step Towards Understanding Jailbreaking VLMs
cs.CR
Submitted: 2025-11-11
Updated: 2026-09-06
Comments: Accepted to EMNLP 2026. Revised version with additional experiments and analysis
Code: https://github.com/Yuxuan2003/weak-ood-jailbreak
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large Vision-Language Models (VLMs) are susceptible to jailbreak attacks: researchers have developed various attack strategies that bypass the safety mechanisms of VLMs.
Terminology
Abstract
Large Vision-Language Models (VLMs) are susceptible to jailbreak attacks: researchers have developed various attack strategies that bypass the safety mechanisms of VLMs. Among these approaches, jailbreak methods based on the Out-of-Distribution (OOD) strategy have garnered widespread attention due to their simplicity and effectiveness. This paper further advances the understanding of OOD-based VLM jailbreak methods. We show that mild OOD manipulations can achieve stronger jailbreak performance than both clean inputs and overly strong perturbations, a non-monotonic pattern we define as "weak-OOD". We explain this phenomenon through a trade-off between two dominant factors: input intent perception and model refusal triggering. Our evidence suggests that these two factors respond differently to OOD manipulations, which is consistent with a discrepancy between broad pre-training robustness and narrower safety alignment. Building on this insight, we draw inspiration from optical character recognition (OCR) capability enhancement---a core task in the pre-training phase of mainstream VLMs. Leveraging this capability, we design JOCR (Jailbreak via OCR-Aware Embedded Text Perturbation), a practical OCR-readable extension of embedded-text jailbreaks that achieves the best average ASR among evaluated baselines. Code is available at GitHub: https://github.com/Yuxuan2003/weak-ood-jailbreak.
Sources
- GPT-4 Technical Report
- Qwen2.5-VL Technical Report
- Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- GPT-4o System Card
- HiddenDetect: Detecting Jailbreak Attacks against Large Vision-Language Models via Monitoring Hidden States
- JailbreakZoo: Survey, Landscapes, and Horizons in Jailbreaking Large Language and Vision-Language Models
- Scaling Laws for Neural Language Models
- GeoShield: Safeguarding Geolocation Privacy from Vision-Language Models via Adversarial Perturbations
- JailBreakV: A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak Attacks
- Visual-RolePlay: Universal Jailbreak Attack on MultiModal Large Language Models via Role-playing Image Character
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
- Jailbreaking Attack against Multimodal Large Language Model
- Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
- Safety Alignment Should Be Made More Than Just a Few Tokens Deep
- CodeAttack: Revealing Safety Generalization Challenges of Large Language Models via Code Completion
- Failures to Find Transferable Image Jailbreaks Between Vision-Language Models
- ImgTrojan: Jailbreaking Vision-Language Models with ONE Image
- Safety Reasoning with Guidelines
- MIRAGE: Multimodal Immersive Reasoning and Guided Exploration for Red-Team Jailbreak Attacks
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs