NeuronFuzz: Safety Neuron Guided Fuzzing for LLM Safety Evaluation
cs.LG, cs.AI, cs.CR, cs.SE
Submitted: 2026-08-26
Updated: 2026-08-26
Code: https://github.com/alexalbertt/jailbreakchat
License: http://creativecommons.org/licenses/by/4.0/
The gist: Safety evaluation is critical for assessing whether aligned Large Language Models (LLMs) remain robust against jailbreak attacks.
Terminology
Abstract
Safety evaluation is critical for assessing whether aligned Large Language Models (LLMs) remain robust against jailbreak attacks. Existing automated testing methods, however, largely rely on response-level feedback: each candidate prompt typically requires generating a target-model response to evaluate its attack effectiveness. This process is expensive and, more importantly, provides only sparse guidance on strongly aligned models, where most candidates are rejected with the same failure outcome. This paper presents NeuronFuzz, a white-box fuzzing framework that exploits internal safety neurons as continuous execution feedback for LLM safety evaluation. A SafetyOracle converts safety-neuron activations into a continuous safety alarm score that serves as feedback for fuzzing and can be obtained during prefill, eliminating response generation from the fuzzing loop. To construct the SafetyOracle, NeuronFuzz uses template-invariant harmful and benign inputs and stability-aware selection to identify a compact set of safety neurons whose activations capture harmful-intent recognition. Moreover, since the safety alarm score is differentiable, NeuronFuzz uses its gradients to identify safety-sensitive template positions and a masked language model to generate fluent, context-compatible mutations while preserving original harmful payload and avoiding additional optimization variables. We evaluate NeuronFuzz across 21 text and multimodal models. Across five white-box source models, it achieves a 76-100% jailbreak discovery rate, outperforming baselines by up to 48 percentage points. Its optimized templates further transfer zero-shot to open-weight and six proprietary target models, achieving average ASR and top-5 ensemble ASR (EASR) of 69.6%/92.6% and 44.1%/60.0%, respectively.
Sources
- Phi-4 Technical Report
- gpt-oss-120b & gpt-oss-20b Model Card
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- Detecting Language Model Attacks with Perplexity
- Gemma 3 Technical Report
- The Llama 3 Herd of Models
- Safety Alignment Can Be Not Superficial With Explicit Safety Signals
- Jailbreaking ChatGPT via Prompt Engineering: An Empirical Study
- A Single Neuron Is Sufficient to Bypass Safety Alignment in Large Language Models
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
- SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks
- The Jailbreak Tax: How Useful are Your Jailbreak Outputs?
- Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs
- The Linear Representation Hypothesis and the Geometry of Large Language Models
- A StrongREJECT for Empty Jailbreaks
- Gemma 4 Technical Report
- Kimi-VL Technical Report
- GoodVibe: Security-by-Vibe for LLM-Based Code Generation
- Steering Language Models With Activation Engineering
- Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks