Depth, Not Breadth: Best-of-N Jailbreaking Beyond Surface Noise
cs.CR, cs.AI
Submitted: 2026-07-29
Updated: 2026-09-25
Code: https://github.com/vacantfury/imaging_text_
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation
- Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
- Compositional Jailbreaking: An Empirical Analysis of Mutator Chain Interactions in Aligned LLMs
- On Evaluating Adversarial Robustness
- Statistical Estimation of Adversarial Risk in Large Language Models under Best-of-N Sampling
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- SRTJ: Self-Evolving Rule-Driven Training-Free LLM Jailbreaking
- The Great Pretender: A Stochasticity Problem in LLM Jailbreak
- LLM Self Defense: By Self Examination, LLMs Know They Are Being Tricked
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs