FUSE: An Evaluating Framework for Dangerous Capabilities of LLMs
cs.AI
Submitted: 2026-09-02
Updated: 2026-09-02
License: http://creativecommons.org/licenses/by/4.0/
The gist: Fragmented safety evaluation undermines the governance of dangerous AI capabilities.
Terminology
Abstract
Fragmented safety evaluation undermines the governance of dangerous AI capabilities. We present a modular framework that evaluates each model through three orthogonal pipelines---Knowledge (K), Defense (D), and Harm (H)---under a unified protocol, aggregating results into a standardized dangerous-capability profile ϕ. Pluggable modules supply scenario seeds, knowledge banks, hazard queries, and judge rubrics, while the core evaluation engine remains unchanged across domains; the CB evaluation is complemented by a cyber pilot demonstrating protocol transfer. Instantiating the framework with a chemical-biological (CB) module, we evaluate 12 commercial LLMs from four families. Our first contribution is a horizontal comparison of dangerous capability across models and model families: the three dimensions expose sharply divergent profiles---models with comparable knowledge differ in refusal resilience, and strong defenders do not generate less harmful content when they do comply---while family-level patterns further separate Claude, DeepSeek, and GPT models. The second is a temporal analysis of capability evolution: tracking K, D, and H against model release dates reveals that dangerous capability has not monotonically declined; newer models deepen knowledge while only partially improving defense, showing that scaling and alignment progress do not uniformly translate into safety. Reliability is established via cross-judge consistency (bootstrap ρ> 0.79, 4 of 5 judges) and pipeline orthogonality (K -- D -- H inter-correlations ρ in [0.32, 0.52]).
Sources
- Frontier AI Regulation: Managing Emerging Risks to Public Safety
- Constitutional AI: Harmlessness from AI Feedback
- More Agents Is All You Need
- The Foundation Model Transparency Index
- LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet
- The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning
- Jailbreaking Black Box Large Language Models in Twenty Queries
- Holistic Evaluation of Language Models
- AgentBench: Evaluating LLMs as Agents
- SciKnowEval: Evaluating Multi-level Scientific Knowledge of Large Language Models
- AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models
- Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned
- Formalizing and Benchmarking Prompt Injection Attacks and Defenses
- Time Travel in LLMs: Tracing Data Contamination in Large Language Models
- Measuring Massive Multitask Language Understanding
- Training language models to follow instructions with human feedback
- Red Teaming Language Models with Language Models
- DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model
- Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection