Efficient Safety Benchmarking via Item Response Theory
cs.CY, cs.AI
Submitted: 2026-05-26
Updated: 2026-09-25
Code: https://github.com/nd-ball/py-irt
Terminology
Sources
- JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
- Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned
- WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
- Stochastic Variational Inference
- The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
- MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
- SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal
- AIR-Bench 2024: A Safety Benchmark Based on Risk Categories from Regulations and Policies
- DISCO: Diversifying Sample Condensation for Efficient Model Evaluation
- A StrongREJECT for Empty Jailbreaks
- InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents
- SafetyBench: Evaluating the Safety of Large Language Models
- SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models
- Universal and Transferable Adversarial Attacks on Aligned Language Models
- Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs
Related papers
- Reasoning Enhances Robustness to Prompt Injection in LLM-Based Consensus
- Generative AI Purpose-built for Social and Mental Health: A Real-World Pilot
- PersonaMem-v3: Toward Omni-Platform Personal Intelligence for Holistic User Understanding, Recommendation, and Agentic Tasks
- What is an intelligent system?
- AI University: An LLM-Powered Learning Assistant for Engineering---A Finite Element Method Case Study
- Generative AI Use in Entrepreneurship: An Integrative Review and an Empowerment-Entrapment Framework