NeuroBreak: Unveil Internal Jailbreak Mechanisms in Large Language Models
cs.CR, cs.AI
Submitted: 2025-09-04
Updated: 2026-08-29
Terminology
Sources
- Understanding intermediate layers using linear classifier probes
- LoRA: Low-Rank Adaptation of Large Language Models
- A Cross-Language Investigation into Jailbreak Attacks in Large Language Models
- Safety Layers in Aligned Large Language Models: The Key to LLM Security
- AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models
- Quantized Delta Weight Is Safety Keeper
- A Simple and Effective Pruning Approach for Large Language Models
- Universal and Transferable Adversarial Attacks on Aligned Language Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs