A Theory of Conditional Collapse under Low-Rank Weight-Space Ablations: I. The Single-Block Theory and Synthetic Validation
Abdallah Khemais
ISITCOM, University of Sousse
cs.LG, cs.AI
Submitted: 2026-08-07
Updated: 2026-08-10
Comments: 25 pages, 2 figures. Part I of a two-part series; see the companion paper "Cross-Layer Interaction under Weight-Space Ablation" (Part II)
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 63/100
The gist: This paper investigates the relationship between activation patching and weight-space ablation, noting that while both are used to argue that a component is causally responsible for a behavior, "they
Terminology
Summary
This paper investigates the relationship between activation patching and weight-space ablation, noting that while both are used to argue that a component is causally responsible for a behavior, they act on different objects: one forward pass, versus the parameters behind every forward pass.
The study utilizes an idealized model in which a conditional computation is carried additively through a residual stream, F(x) = F 0(x) + sum i=1 k alpha i(x)v i, and read out by a linear functional, s(x) = psi(F(x)) + b.
The paper presents three primary theoretical results:
1. Conditional Collapse
The authors establish the conditions under which deleting a subset of carriers S causes a matched input pair (x A, x B) to collapse to the same output. "Deleting a subset of the carriers v i collapses a matched input pair onto one and the same unconditional output if and only if the removal is symmetric on the pair and leaves no contrast outside it; the resulting error is deterministic, with polarity the sign of the pair’s mean margin." Formally, s S(x A) = s S(x B) = = 1 over 2(g A + g B) if and only if:
-
(i) S = 0 (the removal is symmetric on the pair), and
-
(ii) S = 0 (no contrast survives outside S).
The paper also provides a robust form
(Proposition 1), stating that if S epsilon 1 and S epsilon 2, then... both inputs of the pair are mapped to the branch sign
provided > epsilon 1 + 1 over 2 epsilon 2.
2. Patching–Ablation Dissociation
The paper proves that patching and ablation are governed by different quantities: patching a carrier moves the readout by that carrier’s donor–receiver contrast, whereas ablating it moves the readout by its absolute level at the receiver.
Specifically, patching i into the receiver changes the selector by exactly beta i delta i [...] whereas ablating i alone changes the selector by exactly-beta i alpha i(x B).
Because neither quantity bounds the other,
the authors construct a redundancy regime in which patching overstates importance and ablation understates it,
specifically a regime on which every single-carrier patch flips the decision while no single-carrier ablation does.
3. Nonlinear Interaction
The authors address the breakdown of the independence assumption in the idealized model, specifically for an attention head composed with its own layer’s normalization and MLP.
They derive an exact first-order formula, with a provably second-order remainder, for the interaction term the idealized model sets to zero
:
(x) = (I - Q)Dg(r 1(x)) eta(x) - R(x)
The paper shows that (x) vanishes identically whenever the MLP alone is ablated but not, in general, when a head is.
This interaction term identifies the single architectural mechanism through which the first two results cease to be exact.
Synthetic Validation
The theoretical predictions were tested on small transformers trained on a synthetic conditional task. The results showed:
-
Interaction Correlation:
Over thirty-nine ablation configurations the measured interaction is strongly rank-correlated with how well the idealized model predicts the edited network’s behavior (Spearman-0.83).
-
Monotone Relationship: The authors found that
a clean separation [between configurations that follow the idealized model and those that do not] does not survive... we report the weaker monotone claim that does.
-
Replication: A second task and architecture (an inverse, value-to-key recall)
reproduces the same monotone relationship, the same patch-equals-weight-edit exactness, and a further instance of the polarity reversal.
-
Protocol Accuracy: The
patch-equals-weight-edit
identity (Proposition 2) was confirmed, showing that for a single carrier, theweight-edit route... and the frozen-activation route... assign identical values to every node of the network.
Improvements for AI systems
1. Interaction-Aware Model Editing
-
The Improvement: Integration of the (x) interaction formula into automated weight-editing pipelines. Instead of performing simple weight ablation, the system calculates the predicted second-order remainder caused by the MLP and LayerNorm components.
-
What the improved system can do: It can perform
zero-remainder
edits. When attempting to remove a specific behavior (such as a bias or a factual error), the system will simultaneously adjust the weights of the MLP and LayerNorm to neutralize the nonlinear compensation that usually occurs when an attention head is ablated, ensuring the unwanted behavior is truly eliminated rather than just masked.
2. Dual-Metric Interpretability Diagnostics
-
The Improvement: A diagnostic framework that replaces single-metric importance scores (which currently rely on either patching or ablation) with a dual-metric
Importance Profile.
This profile calculates both the donor-receiver contrast (the patching delta) and the absolute level (the ablation delta) for every component. -
What the improved system can do: It can distinguish between
functional importance
andpresence importance.
This allows researchers to identifyredundant
components that appear critical during activation patching but are actually non-essential, preventing the misallocation of compute or human effort during mechanistic interpretability studies.
3. Contrast-Preserving Pruning Algorithms
-
The Improvement: A pruning objective function based on the
Conditional Collapse
theoretical results. Rather than pruning based on weight magnitude, the algorithm evaluates the S (residual contrast) and S (symmetry) conditions for candidate subsets of components. -
What the improved system can do: It can perform
non-collapsing
compression. The system can prune large numbers of parameters while mathematically guaranteeing that the removal does not trigger aconditional collapse,
ensuring the model retains its ability to distinguish between specific input pairs (x A, x B) and maintains its decision-making margins.
Abstract
Activation patching and weight-space ablation both claim a component is causally responsible for a behavior, yet they act on different objects: one forward pass versus the parameters behind every forward pass. We ask when they agree. We study an idealized model where a conditional computation is carried additively through a residual stream, F(x)=F 0(x)+ sum i alpha i(x)v i, read out by a linear functional, and prove three exact results. First, deleting a subset of carriers collapses a matched input pair onto the same unconditional output if and only if the removal is symmetric on the pair and leaves no outside contrast; the error is deterministic, and we give its exact form even when the two conditions hold only approximately. Second, patching a carrier moves the readout by its donor-receiver contrast, while ablating it moves the readout by its absolute level; neither bounds the other, and we construct pairs where every single-carrier patch flips the decision while no single-carrier ablation does. Third, for an attention head composed with its own layer's normalization and MLP, we derive an exact first-order interaction formula with a provably second-order remainder, vanishing identically when only the MLP is ablated but not, in general, when a head is. Small transformers trained on a synthetic conditional task illustrate all three predictions: across thirty-nine ablation configurations the measured interaction is strongly rank-correlated with the idealized model's predictive accuracy (Spearman-0.83), and a second task and architecture reproduces the same pattern, including a further polarity reversal. The single-block interaction result extends past one residual block, and the synthetic validation is tested against a real pretrained model, in a companion paper that takes this theory further along both axes.
Sources
- The Hydra Effect: Emergent Self-repair in Language Model Computations
- Localizing Model Behavior with Path Patching
- The Curse of Multiple Mediators: Hidden Interaction Effects in Activation Patching
- Conditional Co-Ablation: Recovering Self-Repair Backups in Transformer Circuits
- Beyond Importance: Interchange-Sobol Sensitivity Reveals Task-Specific Content Channels in Transformer Components
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks