The Fragility of Jailbreak Robustness Across Operational States
cs.CR, cs.CL
Submitted: 2026-08-31
Updated: 2026-08-31
Code: https://github.com/patrickrchao/JailbreakingLLMshttps:
Terminology
Sources
- The Llama 3 Herd of Models
- Mistral 7B
- DeepInception: Hypnotize Large Language Model to Be Jailbreaker
- AmpleGCG: Learning a Universal and Transferable Generative Model of Adversarial Suffixes for Jailbreaking Both Open and Closed LLMs
- AdvPrompter: Fast Adaptive Adversarial Prompting for LLMs
- Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Qwen2.5 Technical Report
- The Better Angels of Machine Personality: How Personality Relates to LLM Safety
- BERTScore: Evaluating Text Generation with BERT
- Enhancing Jailbreak Attacks on LLMs via Persona Prompts
- Representation Engineering: A Top-Down Approach to AI Transparency
- Universal and Transferable Adversarial Attacks on Aligned Language Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs