A Theory of Conditional Collapse under Low-Rank Weight-Space Ablations: I. The Single-Block Theory and Synthetic Validation

arXiv:2608.03620 · cs.LG, cs.AI · Submitted 2026-08-07 · Read on arXiv

Abdallah Khemais

ISITCOM, University of Sousse

cs.LG, cs.AI

Submitted: 2026-08-07

Updated: 2026-08-10

Comments: 25 pages, 2 figures. Part I of a two-part series; see the companion paper "Cross-Layer Interaction under Weight-Space Ablation" (Part II)

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 63/100

The gist: This paper investigates the relationship between activation patching and weight-space ablation, noting that while both are used to argue that a component is causally responsible for a behavior, "they

Terminology

Summary

This paper investigates the relationship between activation patching and weight-space ablation, noting that while both are used to argue that a component is causally responsible for a behavior, they act on different objects: one forward pass, versus the parameters behind every forward pass. The study utilizes an idealized model in which a conditional computation is carried additively through a residual stream, F(x) = F 0(x) + sum i=1 k alpha i(x)v i, and read out by a linear functional, s(x) = psi(F(x)) + b.

The paper presents three primary theoretical results:

1. Conditional Collapse

The authors establish the conditions under which deleting a subset of carriers S causes a matched input pair (x A, x B) to collapse to the same output. "Deleting a subset of the carriers v i collapses a matched input pair onto one and the same unconditional output if and only if the removal is symmetric on the pair and leaves no contrast outside it; the resulting error is deterministic, with polarity the sign of the pair’s mean margin." Formally, s S(x A) = s S(x B) = = 1 over 2(g A + g B) if and only if:

  • (i) S = 0 (the removal is symmetric on the pair), and

  • (ii) S = 0 (no contrast survives outside S).

The paper also provides a robust form (Proposition 1), stating that if S epsilon 1 and S epsilon 2, then... both inputs of the pair are mapped to the branch sign provided > epsilon 1 + 1 over 2 epsilon 2.

2. Patching–Ablation Dissociation

The paper proves that patching and ablation are governed by different quantities: patching a carrier moves the readout by that carrier’s donor–receiver contrast, whereas ablating it moves the readout by its absolute level at the receiver. Specifically, patching i into the receiver changes the selector by exactly beta i delta i [...] whereas ablating i alone changes the selector by exactly-beta i alpha i(x B). Because neither quantity bounds the other, the authors construct a redundancy regime in which patching overstates importance and ablation understates it, specifically a regime on which every single-carrier patch flips the decision while no single-carrier ablation does.

3. Nonlinear Interaction

The authors address the breakdown of the independence assumption in the idealized model, specifically for an attention head composed with its own layer’s normalization and MLP. They derive an exact first-order formula, with a provably second-order remainder, for the interaction term the idealized model sets to zero:

(x) = (I - Q)Dg(r 1(x)) eta(x) - R(x)

The paper shows that (x) vanishes identically whenever the MLP alone is ablated but not, in general, when a head is. This interaction term identifies the single architectural mechanism through which the first two results cease to be exact.

Synthetic Validation

The theoretical predictions were tested on small transformers trained on a synthetic conditional task. The results showed:

  • Interaction Correlation: Over thirty-nine ablation configurations the measured interaction is strongly rank-correlated with how well the idealized model predicts the edited network’s behavior (Spearman-0.83).

  • Monotone Relationship: The authors found that a clean separation [between configurations that follow the idealized model and those that do not] does not survive... we report the weaker monotone claim that does.

  • Replication: A second task and architecture (an inverse, value-to-key recall) reproduces the same monotone relationship, the same patch-equals-weight-edit exactness, and a further instance of the polarity reversal.

  • Protocol Accuracy: The patch-equals-weight-edit identity (Proposition 2) was confirmed, showing that for a single carrier, the weight-edit route... and the frozen-activation route... assign identical values to every node of the network.

Improvements for AI systems

1. Interaction-Aware Model Editing

  • The Improvement: Integration of the (x) interaction formula into automated weight-editing pipelines. Instead of performing simple weight ablation, the system calculates the predicted second-order remainder caused by the MLP and LayerNorm components.

  • What the improved system can do: It can perform zero-remainder edits. When attempting to remove a specific behavior (such as a bias or a factual error), the system will simultaneously adjust the weights of the MLP and LayerNorm to neutralize the nonlinear compensation that usually occurs when an attention head is ablated, ensuring the unwanted behavior is truly eliminated rather than just masked.

2. Dual-Metric Interpretability Diagnostics

  • The Improvement: A diagnostic framework that replaces single-metric importance scores (which currently rely on either patching or ablation) with a dual-metric Importance Profile. This profile calculates both the donor-receiver contrast (the patching delta) and the absolute level (the ablation delta) for every component.

  • What the improved system can do: It can distinguish between functional importance and presence importance. This allows researchers to identify redundant components that appear critical during activation patching but are actually non-essential, preventing the misallocation of compute or human effort during mechanistic interpretability studies.

3. Contrast-Preserving Pruning Algorithms

  • The Improvement: A pruning objective function based on the Conditional Collapse theoretical results. Rather than pruning based on weight magnitude, the algorithm evaluates the S (residual contrast) and S (symmetry) conditions for candidate subsets of components.

  • What the improved system can do: It can perform non-collapsing compression. The system can prune large numbers of parameters while mathematically guaranteeing that the removal does not trigger a conditional collapse, ensuring the model retains its ability to distinguish between specific input pairs (x A, x B) and maintains its decision-making margins.

Abstract

Activation patching and weight-space ablation both claim a component is causally responsible for a behavior, yet they act on different objects: one forward pass versus the parameters behind every forward pass. We ask when they agree. We study an idealized model where a conditional computation is carried additively through a residual stream, F(x)=F 0(x)+ sum i alpha i(x)v i, read out by a linear functional, and prove three exact results. First, deleting a subset of carriers collapses a matched input pair onto the same unconditional output if and only if the removal is symmetric on the pair and leaves no outside contrast; the error is deterministic, and we give its exact form even when the two conditions hold only approximately. Second, patching a carrier moves the readout by its donor-receiver contrast, while ablating it moves the readout by its absolute level; neither bounds the other, and we construct pairs where every single-carrier patch flips the decision while no single-carrier ablation does. Third, for an attention head composed with its own layer's normalization and MLP, we derive an exact first-order interaction formula with a provably second-order remainder, vanishing identically when only the MLP is ablated but not, in general, when a head is. Small transformers trained on a synthetic conditional task illustrate all three predictions: across thirty-nine ablation configurations the measured interaction is strongly rank-correlated with the idealized model's predictive accuracy (Spearman-0.83), and a second task and architecture reproduces the same pattern, including a further polarity reversal. The single-block interaction result extends past one residual block, and the synthetic validation is tested against a real pretrained model, in a companion paper that takes this theory further along both axes.

Sources

Related papers