Learning to Detect Unseen Jailbreak Attacks in Large Vision-Language Models
cs.CR, cs.AI, cs.CV
Submitted: 2025-08-08
Updated: 2026-08-26
Code: https://github.com/meta-llama/PurpleLlama
Terminology
Sources
- Detecting Language Model Attacks with Perplexity
- Qwen2.5-VL Technical Report
- Why Should Adversarial Perturbations be Imperceptible? Rethink the Research Paradigm in Adversarial NLP
- A Comparative Study of Rule-Based and Data-Driven Approaches in Industrial Monitoring
- MirrorCheck: Efficient Adversarial Defense for Vision-Language Models
- FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts
- HiddenDetect: Detecting Jailbreak Attacks against Large Vision-Language Models via Monitoring Hidden States
- JailbreakZoo: Survey, Landscapes, and Horizons in Jailbreaking Large Language and Vision-Language Models
- A Survey of Attacks on Large Vision-Language Models: Resources, Advances, and Future Trends
- Safety of Multimodal Large Language Models on Images and Texts
- JailBreakV: A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak Attacks
- Visual-RolePlay: Universal Jailbreak Attack on MultiModal Large Language Models via Role-playing Image Character
- Jailbreaking Attack against Multimodal Large Language Model
- MLLM-Protector: Ensuring MLLM's Safety without Hurting Performance
- Jailbreak Large Vision-Language Models Through Multi-Modal Linkage
- GradSafe: Detecting Jailbreak Prompts for LLMs via Safety-Critical Gradient Analysis
- Cross-modality Information Check for Detecting Jailbreaking in Multimodal Large Language Models
- A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations
- MM-Vet v2: A Challenging Benchmark to Evaluate Large Multimodal Models for Integrated Capabilities
- FC-Attack: Jailbreaking Multimodal Large Language Models via Auto-Generated Flowcharts
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs