Cross-Layer Interaction under Weight-Space Ablation: A Closed-Form Attention Jacobian Bound and a Test on a Real Pretrained Model
Abdallah Khemais
ISITCOM, University of Sousse
cs.AI, cs.LG
Submitted: 2026-08-07
Updated: 2026-08-10
Comments: 18 pages, 2 figures. Part II of a two-part series; see the companion paper "A Theory of Conditional Collapse under Low-Rank Weight-Space Ablations" (Part I)
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 54/100
The gist: This paper extends the study of interaction effects between different components of a transformer model from a single residual block to a multi-layer setting and tests these findings on a real
Terminology
Summary
This paper extends the study of interaction effects between different components of a transformer model from a single residual block to a multi-layer setting and tests these findings on a real pretrained model. The research addresses two primary gaps left by a companion paper: the limitation of the single-block interaction theorem to a single residual block and the reliance on small, synthetic transformers for empirical verification.
Multi-Layer Decomposition
The paper first addresses the interaction produced by ablating an arbitrary subset of carriers spanning multiple layers. It establishes that "the gap between the true, weight-edited network’s selector and the idealized model’s own prediction decomposes exactly into a sum of same-block terms, one per touched layer, each pinned to zero when that layer’s own carriers are MLP-only and each equal to the companion paper’s bounded interaction term when the layer feeds the readout directly, plus a cross-layer remainder on which the decomposition itself makes no claim of smallness." This is formalized in Proposition 1, which provides an exact decomposition for any matched pair and any ablated subset S.
Cross-Layer Interaction Identity and Attention Jacobian Bound
The paper further isolates the cross-layer remainder for the case of two touched layers. It defines this remainder as an exact identity: a double integral of a mixed second derivative, the same first-order technique applied one level up.
To close this identity into a usable numerical bound, the paper identifies the need for a Jacobian bound for the attention sub-block itself.
The authors derive that bound in closed form and verify it, pointwise and without a single violation, against Qwen2.5-1.5B-Instruct’s real weights.
However, the paper notes that it does not yet chain it across many layers,
stating that while the per-layer bound is verified, the chained, multi-layer one [is] unverified and likely loose.
Closed-Form Curvature Constant
The paper also addresses the curvature constant used in the companion paper’s second-order interaction bound. Rather than assuming the existence of this constant, the authors compute in closed form, in terms of the three MLP spectral norms and the normalization scale of the trained weights alone, discharging that hypothesis rather than assuming it.
This is detailed in Proposition 4, which provides a bound for D squared g(r) op that is multiplied by rho(r) squared rather than growing with r 1(x),
ensuring the remainder does not blow up simply because the residual stream does.
Empirical Test on a Real Pretrained Model
Finally, the paper conducts an empirical test on Qwen2.5-1.5B-Instruct to see if the qualitative predictions of the companion paper hold for an emergent circuit in a genuinely pretrained model. Using the original activation-patching method, the authors find an emergent circuit for indirect object identification, a mechanism nobody trained the model to have and nobody designed into it.
The results of testing this circuit for the three-way pattern of collapse, dissociation, and interaction are described as genuinely mixed
:
-
Shared Circuit: A
small shared core plus instances-specific idiosyncratic support
was found, with one head appearing in all five tested instances. -
Collapse and Dissociation: The results show that
collapse and dissociation are present on most instances and absent or weak on at least one each.
-
Interaction: A "nonzero interaction is measurable on three of the five, at layer pairs that sit outside the same-block configuration the companion theorem was proven for, so what they show bears on this paper’s own open multi-layer case rather than on that theorem directly."
Improvements for AI systems
Surgical Weight-Editing for Safety Alignment
- What it can do: Uses the multi-layer decomposition formula to perform highly localized weight updates that neutralize specific
emergent circuits
(e.g., harmful biases or misinformation) while providing a mathematical guarantee that thecross-layer remainder
remains below a specific threshold. This prevents theside-effect
problem where fixing one behavior inadvertently degrades unrelated capabilities in other layers.
Error-Bounded Model Compression and Pruning
- What it can do: Employs the closed-form curvature constant and the derived Jacobian bounds to identify and prune specific attention heads or MLP blocks. The system can automatically shrink a model's parameter count while guaranteeing that the total performance gap—calculated as the sum of same-block terms and the cross-layer remainder—does not exceed a predefined error budget.
Automated Mechanistic Debugging Tools
- What it can do: Utilizes the cross-layer interaction identity to automatically map out
failure circuits
in real-world models. When a model fails a specific reasoning task (such as indirect object identification), the system can pinpoint the exact layer-pair interaction responsible for the error, allowing developers to perform targeted fine-tuning on specific functional nodes rather than retraining the entire model.
Stability-Certified Training Optimizers
- What it can do: Integrates the closed-form (calculated via MLP spectral norms and normalization scales) into the training loop to monitor second-order curvature. This allows for an optimizer that automatically adjusts learning rates or weight scales to ensure the residual stream does not cause the error remainder to
blow up,
resulting in more stable training for ultra-deep transformer architectures.
Interaction-Aware Architecture Search (NAS)
- What it can do: Uses the mathematical bounds on attention-sub-block Jacobians to guide the design of new transformer architectures. Instead of simply adding more layers, the system optimizes the arrangement of MLP and attention blocks to minimize the
cross-layer remainder,
creating more efficient models that achieve higher accuracy with fewer total layers.
Abstract
A companion paper studies when activation patching and weight-space ablation agree, inside an idealized model where a conditional computation is carried additively through a residual stream. For the one composition in that model where two carriers are architecturally dependent, an attention head and its own layer's normalization-MLP composition, it derives an exact first-order interaction formula, zero when only the MLP is ablated and second-order bounded when the head is also ablated. That result is confined to a single residual block and checked only on small transformers on a synthetic task. This paper extends the result past both limits. First, the interaction from ablating carriers spanning several layers decomposes exactly into same-block terms, one per touched layer, plus a cross-layer remainder on which the decomposition makes no claim of smallness. Second, we isolate that remainder exactly, for two layers, as a double integral of a mixed second derivative, and name the missing ingredient needed to bound it: a Jacobian bound for the attention sub-block. We derive this bound in closed form and verify it, without a single violation, against Qwen2.5-1.5B-Instruct's real weights, though we do not yet chain it across layers. We also give, in closed form, the curvature constant the companion paper's bound leaves unexhibited. Third, on that same model, we search for and find an emergent circuit for indirect object identification, never designed into it, using the original activation-patching method for this task, and test collapse, dissociation, and interaction on it. The result is mixed: a shared carrier emerges across all five tested instances, collapse and dissociation hold on most but not all, and a nonzero interaction is measurable on three of five, at layer pairs outside the same-block case the companion theorem covers.
Sources
- The Hydra Effect: Emergent Self-repair in Language Model Computations
- The Curse of Multiple Mediators: Hidden Interaction Effects in Activation Patching
- Conditional Co-Ablation: Recovering Self-Repair Backups in Transformer Circuits
- Beyond Importance: Interchange-Sobol Sensitivity Reveals Task-Specific Content Channels in Transformer Components
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection