Voices Across Registers: Corpus-Conditioned Vernacular Jailbreaks against Aligned LLMs via Fanfiction Subgenres
cs.CL
Submitted: 2026-06-03
Updated: 2026-08-30
Comments: Accepted by EMNLP 2026 Main Conference, 25 pages
Code: https://github.com/T-Lab-CUHKSZ/VAR
License: http://creativecommons.org/licenses/by/4.0/
The gist: Existing jailbreaks against aligned LLMs are discrete artifacts whose surface forms are easy to fingerprint and patch.
Terminology
Abstract
Existing jailbreaks against aligned LLMs are discrete artifacts whose surface forms are easy to fingerprint and patch. We argue that the broader failure mode may lie not in any specific prompt, but in natural writing registers that safety tuning under-covers. Building on this insight, we introduce VAR, the first jailbreak family that uses real fanfiction subgenres as universal attack carriers: a creative-writing meta prompt is conditioned on passages from one of twelve Archive of Our Own (AO3) subgenres, and the harmful behavior is embedded as the climax of the resulting scene. The construction requires neither an adversarial attacker LLM nor optimization. On eight aligned LLMs over the union of HarmBench and JailbreakBench, this attack lifts mean ASR from 0.278 to 0.731 under a four-judge ensemble; a factorial decomposition shows the gain is carried by register rather than length or structure. Two active defenses widen rather than narrow the vernacular-to-baseline ratio, indicating that template-targeting defenses merely steer attackers toward register-based attacks like ours. We also propose VAR-A4, a static four-turn extension that attains a mean ASR of 0.924, substantially exceeding three existing multi-turn methods. Our code and data are safely open-sourced at https://github.com/T-Lab-CUHKSZ/VAR.
Sources
- Adversarial Poetry as a Universal Single-Turn Jailbreak Mechanism in Large Language Models
- Chain-of-Lure: A Universal Jailbreak Attack Framework using Unconstrained Synthetic Narratives
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- The Llama 3 Herd of Models
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- DeepInception: Hypnotize Large Language Model to Be Jailbreaker
- Voice Jailbreak Attacks Against GPT-4o
- Qwen3 Technical Report
- Universal and Transferable Adversarial Attacks on Aligned Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering