JailMeter: An Evidence-Based Evaluation Framework for Jailbreak Attacks on Large Language Models
Qingjia Huang, Jingyu Zhang, Jianguo Wu, Yakai Li, Weijuan Zhang, Yankai Rong, Junyi Yao, Shengzhi Zhang, Xiaoqi Jia
cs.CR, cs.AI, cs.CL
Submitted: 2026-07-20
Code: https://github.com/Magi2B0y/JailMeter
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- JADES: A Universal Framework for Jailbreak Assessment via Decompositional Scoring
- Distilling the Knowledge in a Neural Network
- SoK: Systematizing LLM Prompt Security: Taxonomies, Datasets, and Unified Evaluation of Attacks and Defenses
- GuidedBench: Measuring and Mitigating the Evaluation Discrepancies of In-the-wild LLM Jailbreak Methods
- Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL
- Dialogue Injection Attack: Jailbreaking LLMs through Context Manipulation
- Use of LLMs for Illicit Purposes: Threats, Prevention Measures, and Vulnerabilities
- X-Teaming: Multi-Turn Jailbreaks and Defenses with Adaptive Multi-Agents
- LLMs know their vulnerabilities: Uncover Safety Gaps through Natural Distribution Shifts
- GRAF: Multi-turn Jailbreaking via Global Refinement and Active Fabrication
- LLaMA: Open and Efficient Foundation Language Models
- SoK: Evaluating Jailbreak Guardrails for Large Language Models
- Foot-In-The-Door: A Multi-turn Jailbreak for LLMs
- Jigsaw Puzzles: Splitting Harmful Questions to Jailbreak Large Language Models
- Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models
- GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts
- Universal and Transferable Adversarial Attacks on Aligned Language Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs