MechSparse: Mechanism-Guided Sparse PEFT Selection Is Task-Shaped

arXiv:2609.18961 · cs.CL · Submitted 2026-07-16 · Read on arXiv

cs.CL

Submitted: 2026-07-16

Updated: 2026-07-16

License: http://creativecommons.org/licenses/by/4.0/

The gist: Mechanistic interpretability identifies sparse subsets of heads and MLP blocks that carry specific behaviors.

Terminology

Abstract

Mechanistic interpretability identifies sparse subsets of heads and MLP blocks that carry specific behaviors. We ask whether such causal signals can guide where to place a small PEFT budget more effectively than the cheap heuristics practitioners already use. scores attention heads and MLP blocks by normalized activation-patching recovery on clean/corrupted probes and trains LoRA/QLoRA only on the selected sites; adds bounded credit for small within-layer joint subsets. We compare against random, magnitude, activation-norm, and gradient/Fisher on Ministral-8B/NF4 in three cells: Swahili span-JSON information extraction (IE) at b = 0.25% and 1.0%, and English to Swahili machine translation (MT) at b = 1.0%. The causal selectors never win the primary metric. On the headline IE cell (3 seeds, paired-bootstrap CIs over 600 predictions), beats random by +0.079 span+type F1 and gradient/Fisher by +0.174, but trails activation-norm by 0.028, with the smallest cross-seed std (plus or minus 0.003). On MT all four selectors lie within 0.30 BLEU and every paired CI includes zero. A schema-versus-span decomposition explains the IE gap: activation-norm captures the rigid JSON routine, while causal scores track content-sensitive sites. We distill a preliminary diagnostic -- prefer activation-norm when output structure dominates, treat causal selectors as a hypothesis for content-dominated tasks -- and release masks, scores, predictions, and evaluation files for direct replay.

Related papers