GuidedBench: Measuring and Mitigating the Evaluation Discrepancies of In-the-wild LLM Jailbreak Methods
cs.CL, cs.CR
Submitted: 2025-02-24
Updated: 2026-09-01
Comments: Accepted by ICLR 2026; Homepage: https://sproutnan.github.io/GuidedBench/
Project page: https://sproutnan.github.io/GuidedBench
License: http://creativecommons.org/licenses/by/4.0/
The gist: Despite the growing interest in jailbreaks as an effective red-teaming tool for building safe and responsible large language models (LLMs), flawed evaluation system designs have led to significant
Terminology
Abstract
Despite the growing interest in jailbreaks as an effective red-teaming tool for building safe and responsible large language models (LLMs), flawed evaluation system designs have led to significant discrepancies in their effectiveness assessments. With a systematic measurement study based on 37 jailbreak studies since 2022, we find that existing evaluation systems lack case-specific criteria, resulting in misleading conclusions about their effectiveness and safety implications. In this paper, we introduce GuidedBench, a novel benchmark comprising a curated harmful question dataset and GuidedEval, an evaluation system integrated with detailed case-by-case evaluation guidelines. Experiments demonstrate that GuidedBench offers more accurate evaluations of jailbreak performance, enabling meaningful comparisons across methods. GuidedEval reduces inter-evaluator variance by at least 76.03%, ensuring reliable and reproducible evaluations. We reveal why existing jailbreak benchmarks fail to evaluate accurately and suggest better evaluation practices.
Sources
- Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
- JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
- Jailbreaking Black Box Large Language Models in Twenty Queries
- GUARD: Role-playing to Generate Natural-language Jailbreakings to Test Guideline Adherence of Large Language Models
- JailbreakZoo: Survey, Landscapes, and Horizons in Jailbreaking Large Language and Vision-Language Models
- DeepSeek-V3 Technical Report
- SAGE-RT: Synthetic Alignment data Generation for Safety Evaluation and Red Teaming
- Open Sesame! Universal Black Box Jailbreaking of Large Language Models
- From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judge
- Revisiting Jailbreaking for Large Language Models: A Representation Engineering Perspective
- ClassEval: A Manually-Crafted Benchmark for Evaluating LLMs on Class-level Code Generation
- Semantic Mirror Jailbreak: Genetic Algorithm Based Jailbreak Prompts Against Open-source LLMs
- Training Verifiers to Solve Math Word Problems
- Evaluating Large Language Models Trained on Code
- DeepInception: Hypnotize Large Language Model to Be Jailbreaker
- A Survey on LLM-as-a-Judge
- AmpleGCG: Learning a Universal and Transferable Generative Model of Adversarial Suffixes for Jailbreaking Both Open and Closed LLMs
- When "Competency" in Reasoning Opens the Door to Vulnerability: Jailbreaking LLMs via Novel Complex Ciphers
- Goal-Oriented Prompt Attack and Safety Evaluation for LLMs
- Catastrophic Jailbreak of Open-source LLMs via Exploiting Generation
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering