Blind, Not Weak: A Best-of-Suite Safety-Utility Frontier for Recover-and-Reguard Defenses Against Encoded VLM Jailbreaks
cs.CR, cs.AI, cs.LG
Submitted: 2026-07-29
Updated: 2026-09-25
Code: https://github.com/vacantfury/llm_guardrail_security
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Jailbreaking Large Language Models with Symbolic Mathematics
- On Evaluating Adversarial Robustness
- SafePTR: Token-Level Jailbreak Defense in Multimodal LLMs via Prune-then-Restore Mechanism
- Text is All You Need for Vision-Language Model Jailbreaking
- Jailbreaking LLMs & VLMs: Mechanisms, Evaluation, and Unified Defense
- Llama Guard 3 Vision: Safeguarding Human-AI Image Understanding Conversations
- Toward Universal and Transferable Jailbreak Attacks on Vision-Language Models
- Crafting Adversarial Inputs for Large Vision-Language Models Using Black-Box Optimization
- Resolution Thresholds in VLM Detection of Harmful ASCII Art Across Construction Modes and Languages
- Obscure but Effective: Classical Chinese Jailbreak Prompt Optimization via Bio-Inspired Search
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- Cross-Modal Obfuscation for Jailbreak Attacks on Large Vision-Language Models
- The Great Contradiction Showdown: How Jailbreak and Stealth Wrestle in Vision-Language Models?
- Making MLLMs Blind: Adversarial Smuggling Attacks in MLLM Content Moderation
- Jailbreak Attacks and Defenses against Multimodal Generative Models: A Survey
- GuardReasoner-VL: Safeguarding VLMs via Reinforced Reasoning
- Jailbreaking Attack against Multimodal Large Language Model
- Agentic Moderation: Multi-Agent Design for Safer Vision-Language Models
- Seeing the Threat: Vulnerabilities in Vision-Language Models to Adversarial Attack
- Leave My Images Alone: Preventing Multi-Modal Large Language Models from Analyzing Images via Visual Prompt Injection
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs