Do LLMs Know Their Vulnerable Scenarios?
Ziheng Peng, Huiqi Deng, Haoran Jing, Xuankun Rong, Jiahui Han, Xiting Wang, Na Zou, Xia Hu
cs.AI, cs.CR
Submitted: 2026-07-26
Comments: 19 pages, 11 Figures, Under Review
License: http://creativecommons.org/licenses/by/4.0/
The gist: Safety-aligned large language models are trained to refuse harmful requests, yet embedding the same requests in particular scenarios can bypass their safeguards.
Terminology
Abstract
Safety-aligned large language models are trained to refuse harmful requests, yet embedding the same requests in particular scenarios can bypass their safeguards. Existing red-teaming methods empirically identify effective scenarios through observed attack outcomes, but why particular scenarios weaken refusal remains mechanistically unclear. Meanwhile, mechanistic interpretability studies have characterized both refusal directions and jailbreak-associated features, without explaining the relationship between the two representations. In this work, we show that scenario-wrapped prompts activate internal scenario directions whose causal steering consistently reduces refusal scores. Building on this finding, we propose Concept2Scenario, a concept-based attribution framework for vulnerable scenario discovery. It instantiates a broad concept space with a sparse autoencoder, attributes refusal suppression to individual concepts, translates the identified concepts into interpretable natural-language scenarios, and identifies synergistic scenario combinations through interaction attribution. Across three open-source models, two safety benchmarks, and six black-box jailbreak methods, the discovered scenarios serve as reusable priors that improve average attack success rates by up to 18.2 percentage points. They also transfer to GPT-5, Claude-Haiku-4.5, and Gemini-3-Flash, suggesting that some scenario-level refusal vulnerabilities are shared across model families. Moreover, the identified combinations outperform their individual constituents and enable iterative attacks to succeed in fewer turns.
Sources
- CodeChameleon: Personalized Encryption Framework for Jailbreaking Large Language Models
- X-Teaming: Multi-Turn Jailbreaks and Defenses with Adaptive Multi-Agents
- SlimPajama-DC: Understanding Data Combinations for LLM Training
- To Call or Not to Call: Diagnosing Intrinsic Over-Calling Bias in LLM Agents
- OpenAI GPT-5 System Card
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Understanding Refusal in Language Models with Sparse Autoencoders
- Fantastic Reasoning Behaviors and Where to Find Them: Unsupervised Discovery of the Reasoning Process
- Universal and Transferable Adversarial Attacks on Aligned Language Models
- GPT-4 Technical Report
- Sparse Autoencoders are Capable LLM Jailbreak Mitigators
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- The Llama 3 Herd of Models
- DeepInception: Hypnotize Large Language Model to Be Jailbreaker
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection