Reading the Gate, Not the Interference: Output-Side Interference Measurement Does Not Track Merge Collapse

arXiv:2608.11797 · cs.LG · Submitted 2026-08-12 · Read on arXiv

UNSW Sydney

cs.LG

Submitted: 2026-08-12

Updated: 2026-09-01

Code: https://github.com/tatsu-lab/stanford_alpaca

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 75/100

The gist: This paper investigates the causal structure of task-vector interference in merged language models, challenging the prevailing assumption that interference magnitude is the key diagnostic axis.

Terminology

Summary

This paper investigates the causal structure of task-vector interference in merged language models, challenging the prevailing assumption that interference magnitude is the key diagnostic axis. The authors state their central thesis: "in merged language models, interference is transported and amplified along a direction that propagation actively maintains and the prompt format gates at expression—magnitude, local or cumulative, is at best a coarse correlate and is inverted exactly where evaluations look."

The authors track "the exact non-additive residue of merging—the layerwise cross-term Il = hABl − hA0l − h0Bl + h00l—through every block of a merged LLM, measure what each block generates versus what it carries, and then erase the carried term surgically to see what the output loses." The setup uses Qwen2.5-1.5B as the base model with six tasks (math, code, instruction, safety, summarization, translation) finetuned as rank-16 LoRA task vectors, with three training seeds each. The intervention replaces the merged model's residual state at a frozen boundary with hABl − λIl, where λ=1 exactly erases the carried cross-term.

1. Transport dominates generation. An exact decomposition Il+1 = Gl + Tl + Ml (transport of the existing term vs. newly generated pieces) attributes ∼70% of the flux to transport with late-network gain 1.08/block. The authors note: transport carries ∼69% of the layerwise flux (regeneration share 0.31, 3/3 seeds), late-network transport exceeds regeneration by 2.3×, and transport alone amplifies—∥Tl∥/∥Il∥ = 1.08 per late block. A replication on Llama-3.2-1B reproduces both: transport share 65%, late gain 1.14, 3/3 seeds—transport dominance is a family invariant.

2. The carried direction is an attractor. A preregistered regrowth experiment shows erasure is undone by propagation: erase at depth 3 or 8 and the network rebuilds the cross-term essentially exactly—final-boundary ρ = 1.00/0.99, direction cosine to the original cross-term 0.99. Regrowth only falls off when depth runs out: ρ = 0.94 erasing at 14, 0.84 at 20. A basin test with six starting displacements, including 2× overshoot and four structured wrong directions, all reconverging at cosine 0.83–0.99 establishes the carried direction as an attractor of the forward pass.

3. Orientation is causally load-bearing. Erasing along the cross-term's own direction removes expressed interference dose-dependently: Dcorrect = +0.0301/+0.0289/+0.0160 across seeds—removing roughly 19% of expressed interference. The dose-response shows saturation: 0 → +0.017 → +0.025 → +0.026 → +0.020 for λ = 0, 0.5, 1, 1.5, 2 along the correct direction, while the norm-matched cross-prompt direction worsens monotonically: 0 → −0.001 → −0.012 → −0.034 → −0.066. All four norm-matched structural controls fail or backfire, with wrong-pair and coefficient-mismatch patches having negative D.

4. Magnitude is not the signal. The paper's central quantitative contrast: task pairs whose local cross-term generation differs by at most 1.9× differ by 14×–337× in causally removable interference. Specifically, "on the very code prompts where the two merges' local generation differs by at most 1.9× in float32, erasure removes substantial interference from code+safety and almost nothing from code+math—D = +0.0022/+0.0001/+0.0005 across seeds, giving pair ratios of 13.9/337/29.5."

5. Expression gating by instruction wrappers. The authors found that instruction wrappers amplify internal cross-term generation (3/3 seeds, a preregistered prediction falsified in reverse) while collapsing its expression 20×. The identical λ=1 erasure yields "Draw = +0.0286/+0.0370/+0.0174 across seeds but Dwrapped = +0.0031/+0.0026/+0.0005—a 13× mean ratio, raw>wrapped 3/3. The mechanism is denominator-driven: the wrapper inflates the main-effect JSD ∼13× (0.025→0.32) while the absolute interaction JSD does not shrink."

6. A cautionary methodological finding. The striking ±15% cross-condition uniformity of naive bf16 generation turns out to be quantization roughness: "Recomputing Gl on the same common reference state with the block cast to float32 shrinks it to 14–23% of its bfloat16 value, with directions nearly uncorrelated between precisions (median cosine 0.15–0.24). 75–90% of the naively measured 'local generation' is bfloat16 weight-rounding roughness."

The authors report Preregistered reduced replications at a second family (Llama-3.2-1B) and a larger scale (Qwen2.5-7B) reproduce all three pillars. At 7B, "direction-specificity and pair-carrying replicate with all bootstrap intervals excluding zero (pair ratios 6.9–22.9×), the injection exclusion strengthens (up to 8.3× the coherent bound), and gating is total at 7B (Dwrapped indistinguishable from zero, 3/3)."

The authors acknowledge: "(i) Three model points, one adapter parameterization... Full finetuning and further families remain open; the magnitude variation across points is unexplained; we avoid 'universal.' (ii) Intervention statistics initially rested on 3/3 seed-direction agreement alone... (iii) Controlled prompts are diagnostic instruments; gating evidence rests on one wrapper template and two core pairs... (iv) Our preregistered direction features failed on held-out pairs."

The paper concludes: Interference in weight-space composition is not a quantity to be measured but a direction to be traced: transported and amplified, restored when removed, expressed only where the format allows. All 46 predictions were preregistered and frozen before their data, with falsifications including of our own headline expectations, are reported as such.

Improvements for AI systems

Based on this paper, here are specific improvements to AI systems:

Improvement: Replace magnitude-based merging heuristics with direction-tracking algorithms that identify and surgically erase the cross-term attractor direction (Il) rather than minimizing total interference norm.

Capability: When merging multiple fine-tuned models (e.g., combining a math model with a safety model), the system can remove up to 337× more expressed interference on specific task pairs by targeting the carried direction, while leaving task-relevant capabilities intact—something magnitude-based methods fail to achieve.

Improvement: Implement prompt-format gating in merged models—detecting whether the input uses instruction wrappers, raw prompts, or other formats—and dynamically adjust the interference erasure strength (λ) accordingly.

Improvement: Use float32 or mixed-precision computation for the cross-term generation (Gl) during merge-time analysis, filtering out bfloat16 quantization roughness that constitutes 75–90% of naively measured local generation.

Improvement: After erasing interference at an early layer, proactively apply repeated erasure at multiple depths (e.g., layers 3, 8, 14, 20) to counteract the network's tendency to rebuild the cross-term (regrowth to cosine 0.99 within a few layers).

Improvement: Select task pairs for merging based on direction-specific causal removability (D values) rather than generation magnitude, using the finding that pairs with similar generation can differ by 14–337× in removable interference.

Improvement: When evaluating merged models under low-precision inference (bf16), apply a correction factor derived from the float32 cross-term computation to avoid misattributing quantization noise as real interference.

Sources

Related papers