Learning to Detect Unseen Jailbreak Attacks in Large Vision-Language Models

arXiv:2508.09201 · cs.CR, cs.AI, cs.CV · Submitted 2025-08-08 · Read on arXiv

cs.CR, cs.AI, cs.CV

Submitted: 2025-08-08

Updated: 2026-08-26

Code: https://github.com/meta-llama/PurpleLlama

Terminology

Sources

Related papers