FENCE: A Financial and Multimodal Jailbreak Detection Dataset
cs.CL, cs.AI, cs.DB
Submitted: 2026-02-20
Updated: 2026-08-28
Code: https://github.com/kakaobank/FENCE
Terminology
Sources
- GPT-4 Technical Report
- SemEval-2017 Task 1: Semantic Textual Similarity - Multilingual and Cross-lingual Focused Evaluation
- Llama Guard 3 Vision: Safeguarding Human-AI Image Understanding Conversations
- FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts
- A Survey of Attacks on Large Vision-Language Models: Resources, Advances, and Future Trends
- Prompt Injection attack against LLM-integrated Applications
- Jailbreaking ChatGPT via Prompt Engineering: An Empirical Study
- UniGuard: Towards Universal Safety Guardrails for Jailbreak Attacks on Multimodal Large Language Models
- Survey of Vulnerabilities in Large Language Models Revealed by Adversarial Attacks
- PaliGemma 2: A Family of Versatile VLMs for Transfer
- Gemini: A Family of Highly Capable Multimodal Models
- Preventing Jailbreak Prompts as Malicious Tools for Cybercriminals: A Cyber Defense Perspective
- Cross-modality Information Check for Detecting Jailbreaking in Multimodal Large Language Models
- JailGuard: A Universal Detection Framework for LLM Prompt-based Attacks
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering