Localizing Safety Alignment: MLP Layers and Mid-Network Blocks Encode Refusal Behavior in Large Language Models

arXiv:2608.11583 · cs.AI · Submitted 2026-08-12 · Read on arXiv

Mingyu Zong, Sampad Mohanty, Bhaskar Krishnamachari

University of Southern California

cs.AI

Submitted: 2026-08-12

Updated: 2026-08-13

Code: https://github.com/ANRGUSC/Localizing-Safety-Alignment

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: This paper investigates where safety-aligned refusal behavior is encoded in large language models by transplanting weights from aligned models into matched unaligned base models at multiple levels of

Terminology

Summary

This paper investigates where safety-aligned refusal behavior is encoded in large language models by transplanting weights from aligned models into matched unaligned base models at multiple levels of granularity. The authors hypothesize that safety-aligned refusal is localized: the parameters mediating transferable refusal behavior are concentrated in specific weight matrices and depth ranges rather than uniformly distributed across the network.

The study uses two matched open-weight model pairs that share the same architecture within each pair: RealSafe-R1-7B/DeepSeek-R1-Distill-Qwen-7B (28-layer Qwen2-based models) and saferlhf ultra sft/Llama-3.1-8B (32-layer Llama models). The aligned models underwent safety training, while the base models did not. Hybrid models are constructed by replacing selected weight matrices in the base model with corresponding matrices from the aligned model.

The study evaluates several transplantation granularities: component-level hybrids (attention projections vs. MLP projections), contiguous layer groups (first5, mid5, last5, first half, second half), and individual MLP blocks (each containing 4 layers, yielding seven blocks for the 28-layer model and eight blocks for the 32-layer model).

Four safety benchmarks are used: TwinPrompt, SGXSTest, AdvBench (malicious prompts), and OR-Bench (benign prompts). Evaluation subsets are constructed by retaining prompts where the safety-aligned model refuses while the corresponding base model complies, keeping up to 30 prompts per condition.

The first major finding is that transferable refusal behavior is mediated predominantly by the MLP pathway rather than the attention pathway. Under MLP weight transplantation, hybrid models from the RealSafe-R1-7B family achieve 19, 17, and 27 refusals on 30-malicious-prompt subsets from TwinPrompt, SGXSTest, and AdvBench, respectively—the highest MR values among all component-level hybrids. Models from the saferlhf ultra sft family also yield best performance with this configuration. When compared with MR results generated by attn hybrids, MLP-only transplantation exceeds by at least 2.7 times across the six comparisons.

However, the same intervention that restores refusal behavior can also import over-refusal. For RealSafe-R1-7B hybrids, full MLP substitution produces 11 benign refusals on TwinPrompt, whereas attention-only substitution produces none. By contrast, saferlhf ultra sft exhibits no TwinPrompt over-refusal behavior under full MLP transplantation.

Within the MLP stack, Block 3 (layers 8–11) is the most consistently important block across benchmarks. On TwinPrompt and AdvBench, Block 3 achieves the highest MR rates for both model families, and on SGXSTest it appears among the top candidates. "The recurrence of the same absolute depth range across the Qwen2-7B and Llama-3.1-8B architectures suggests that, for models of this scale, the parameters mediating transferable refusal behavior appear disproportionately concentrated around layers 8–11."

Block 3 is selected first in all six greedy searches over model-dataset pairs. Within Block 3, exploratory analysis points to layer 8 for the RealSafe-R1-7B family: "Layer 8 produces more malicious-prompt refusals than the other individual layers. Transplanting only this aligned layer of weights to the base DeepSeek-R1-Distill-Qwen-7B model gives near or above half of the MR that the full block achieves."

The greedy trajectories reveal that aligned blocks do not compose additively: in most settings, adding additional aligned blocks does not guarantee increased refusal, and selective MLP-block subsets can outperform full MLP transplantation. The paper identifies a negative marginal effect where adding an aligned block reduces MR. Five of the six trajectories exhibit such effect, but the disruptive block varies across settings.

For RealSafe-R1-7B on TwinPrompt, MR rises from 7 at k = 1 to 24 at k = 6, then falls to 19 at k = 7. The six-block hybrid achieves MR = 24 with BOR = 9, whereas the full seven-block transplant yields MR = 19 with BOR = 11, which is worse on both metrics. For saferlhf ultra sft-SGXSTest, the six-block model achieves MR = 8 with BOR = 1, compared with MR = 4 and BOR = 3 for the full eight-block model.

The greedy orders transferred to OR-Bench vary with the source benchmark used to derive them. "The AdvBench-derived order is generally more conservative and often incurs higher OR-Bench over-refusal rates. By contrast, the order derived from SGXSTest tends to yield lower OR-Bench over-refusal than the order derived from AdvBench. The authors interpret this as evidence of a precision-coverage trade-off rather than a strict improvement."

The authors validate their findings on larger paired subsets containing 100 malicious and 100 benign prompts. The results strongly reproduce the dominance of the MLP pathway: the mlp hybrid refuses 61 prompts compared with only 8 for the attn hybrid. Block 3 still produces the highest malicious-prompt refusal rates among all seven individual blocks, rejecting 13 of 100 malicious prompts. The greedy trajectory shows "malicious-prompt refusals increase from 13 with one block to 29, 42, 52, 54, and a peak of 64 as the block budget increases from k = 1 to k = 6. However, transplanting the final remaining MLP block reduces performance from 64 to 61 refusals at k = 7."

The paper concludes with three consistent empirical patterns: Pathway Concentration: Transferable refusal behavior is mediated predominantly by the MLP pathway rather than the attention pathway; Depth Localization: Within the MLP stack, a specific mid-depth block (Block 3, layers 8–11) consistently emerges as the strongest localized intervention; and Non-Additive Interactions: Refusal-relevant MLP blocks interact non-monotonically.

The authors emphasize that "safety alignment in current LLMs is both localized and interaction-sensitive. Localization helps explain why aligned behavior can be brittle under subsequent fine-tuning or model modification, while the observed interactions show that effective alignment is not simply a matter of transplanting more aligned parameters."

The authors acknowledge several limitations: the analysis covers only two matched model pairs in the 7B-8B regime; filtered subsets limit sizes; and greedy forward selection only approximates the space of useful block combinations. They suggest future work on combinatorial search, sparse optimization, or learned selection policies that explicitly optimize both malicious refusal and benign non-refusal.

Improvements for AI systems

Improvements to AI Systems:

  1. Targeted Safety Patch Injection: Instead of full-model fine-tuning or alignment retraining, AI systems can implement a safety patch by transplanting only the MLP weight matrices from layers 8–11 of a safety-aligned model into an unaligned base model. This achieves up to 27/30 malicious-prompt refusals with minimal computational cost, enabling rapid deployment of safety behavior onto existing base models without full retraining.

  2. Selective Block Composition for Optimal Safety-Capability Trade-off: AI systems can use a greedy selection algorithm to identify the optimal subset of MLP blocks (e.g., 6 of 7 blocks) rather than transplanting all aligned blocks. This avoids the negative marginal effect where adding the final block reduces refusal performance by up to 5 points and increases over-refusal by 2 points. The system can dynamically choose the block budget to maximize malicious-prompt refusal while minimizing benign-prompt over-refusal.

  3. Benchmark-Aware Safety Calibration: AI systems can derive different safety intervention orders from different benchmarks (e.g., SGXSTest vs. AdvBench) and select the appropriate order based on deployment context. For high-precision environments (e.g., customer service), use the SGXSTest-derived order to reduce over-refusal; for high-coverage environments (e.g., content moderation), use the AdvBench-derived order to maximize malicious-prompt rejection, accepting higher over-refusal.

  4. Layer-Specific Safety Auditing: AI systems can audit their own safety alignment by testing individual layer contributions (e.g., layer 8 in Qwen2-7B) to refusal behavior. This allows for identification of critical safety-relevant parameters, enabling targeted monitoring or protection against accidental unalignment during subsequent fine-tuning or model editing.

  5. Non-Additive Interaction-Aware Alignment: AI systems can implement an alignment scheduler that tracks the marginal contribution of each added safety module (e.g., MLP block) and automatically stops adding modules when the marginal effect becomes negative. This prevents performance degradation from over-alignment and ensures the system maintains optimal refusal rates without unnecessary over-refusal.

Capabilities of the Improved AI System:

  • Can be rapidly safety-aligned by patching only 14% of the model's MLP weights (layers 8–11), reducing alignment cost by an estimated 80% compared to full fine-tuning.

  • Achieves malicious-prompt refusal rates of 19–27 out of 30 while keeping benign over-refusal at 0–11, with the ability to tune this trade-off per deployment.

  • Detects and avoids non-monotonic interactions between safety modules, preventing performance drops from adding too many aligned blocks.

  • Provides interpretable safety diagnostics by identifying which specific layers and pathways (MLP vs. attention) are responsible for refusal, enabling targeted debugging of safety failures.

  • Can be deployed in both high-precision (low over-refusal) and high-coverage (high malicious rejection) modes by swapping the benchmark-derived block selection order.

Abstract

Safety alignment in large language models is often treated as a distributed property of the entire network, yet its practical brittleness suggests that refusal behavior may be concentrated in a smaller set of parameters. This work addresses where safety-aligned refusal is encoded by transplanting weights from aligned models into matched unaligned base models at multiple levels of granularity. Using two open-weight model pairs and four safety benchmarks, we conducted experiments to compare the effects of replacing attention weights, MLP weights, contiguous layer regions, and MLP blocks. Across both model families, refusal transfer is dominated by MLP weights: replacing MLP parameters recovers substantially more malicious-prompt refusal than replacing attention parameters, with gains of at least 2.7 times more across benchmarks. Within the MLP stack, refusal-relevant parameters exhibit a consistent mid-network concentration, as the block spanning layers 8-11 is selected first in all six greedy searches over model-dataset pairs. The results also show that the composition of safety-relevant components is non-additive: in five of six greedy trajectories, adding more aligned blocks can reduce refusal performance, and selective block subsets can outperform full MLP transplantation on malicious refusal, benign over-refusal, or both. Finally, greedy orders transferred to OR-Bench vary with the source benchmark used to derive them, indicating a benchmark-dependent precision-coverage trade-off. These results suggest that safety alignment in current LLMs is both localized and interaction-sensitive, offering insight into alignment brittleness and potential avenues for targeted safety interventions.

Sources

Related papers