Mechanistic Interpretability of LLM Jailbreaks via Internal Attribution Graphs
Anupam Wagle, Ifrat Ikhtear Uddin, Chaowei Zhang, Longwei Wang
cs.CR, cs.AI
Submitted: 2026-07-08
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large language models (LLMs) exhibit remarkable capabilities but remain highly vulnerable to adversarial prompts and jailbreak attacks.
Terminology
Abstract
Large language models (LLMs) exhibit remarkable capabilities but remain highly vulnerable to adversarial prompts and jailbreak attacks. Existing approaches primarily analyze these failures through input-output behaviors or attribution methods, offering limited insight into how adversarial perturbations alter the model's internal reasoning. Consequently, the mechanisms underlying unsafe or incorrect behaviors remain poorly understood. We introduce a mechanistic framework for diagnosing LLM vulnerabilities using paired internal computation graphs, which represent prompt-specific inference as structured causal interactions among latent features. By constructing and aligning computation graphs for clean and attacked prompts, we reveal that adversarial attacks induce systematic transformations of internal reasoning, including suppression of safety-relevant components, emergence of attack-specific features, and rerouting of computation paths. Building on this representation, we propose a unified framework that (i) decomposes computation into invariant, suppressed, and emergent structures, (ii) identifies recurring vulnerability motifs associated with failure modes, and (iii) performs causal interventions on nodes, paths, and subgraphs to directly evaluate their contributions to attack success. This enables a transition from descriptive attribution to causal diagnosis of model failures. Experiments across multiple open-source LLMs and diverse adversarial and jailbreak benchmarks demonstrate that structural deviations in internal computation graphs strongly correlate with unsafe behaviors. Furthermore, targeted interventions on identified vulnerability motifs improve model robustness, establishing internal computation graphs as a principled foundation for understanding, diagnosing, and mitigating LLM vulnerabilities.
Sources
- Explaining and Harnessing Adversarial Examples
- Universal and Transferable Adversarial Attacks on Aligned Language Models
- Bridging Interpretability and Robustness Using LIME-Guided Model Refinement
- Dense Cross-Connected Ensemble Convolutional Neural Networks for Enhanced Model Robustness
- GetNetUPAM: Ecologically Informed Nested Cross-Validation and Noise-Robust Attention for Marine Bioacoustic Monitoring
- CoSwin: Convolution Enhanced Hierarchical Shifted Window Attention For Small-Scale Vision
- Intriguing properties of neural networks
- Deep Inside Convolutional Networks: Visualising Image Classification Models and Saliency Maps
- LLaMA: Open and Efficient Foundation Language Models
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
- Baseline Defenses for Adversarial Attacks Against Aligned Language Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs