Jailbreaking in the Haystack
Rishi Rajesh Shah, Chen Henry Wu, Shashwat Saxena, Ziqian Zhong, Alexander Robey, Aditi Raghunathan
cs.CR, cs.AI, cs.CL, cs.LG
Submitted: 2026-08-19
Updated: 2026-08-20
Code: https://github.com/AR-FORUM/NINJA_Attack
Project page: https://ar-forum.github.io/ninjaattackweb
Terminology
Sources
- Jailbreaking Black Box Large Language Models in Twenty Queries
- MasterKey: Automated Jailbreak Across Multiple Large Language Model Chatbots
- SHADE-Arena: Evaluating Sabotage and Monitoring in LLM Agents
- Refusal-Trained LLMs Are Easily Jailbroken As Browser Agents
- AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models
- LongSafety: Evaluating Long-Context Safety of Large Language Models
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
- Tree of Attacks: Jailbreaking Black-Box LLMs Automatically
- Red Teaming the Mind of the Machine: A Systematic Evaluation of Prompt Injection and Jailbreak Vulnerabilities in LLMs
- Great, Now Write an Article About That: The Crescendo Multi-Turn LLM Jailbreak Attack
- PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks
- Cognitive Overload Attack:Prompt Injection for Long Context
- Robustness, Security, Privacy, Explainability, Efficiency, and Usability of Large Language Models for Code
- A Dataset of Open-Domain Question Answering with Multiple-Span Answers
- Magnon dispersion and spin transport in CrCl$_3$ bilayers under different strain-induced magnetic states
- Universal and Transferable Adversarial Attacks on Aligned Language Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs