MechSparse: Mechanism-Guided Sparse PEFT Selection Is Task-Shaped
cs.CL
Submitted: 2026-07-16
Updated: 2026-07-16
License: http://creativecommons.org/licenses/by/4.0/
The gist: Mechanistic interpretability identifies sparse subsets of heads and MLP blocks that carry specific behaviors.
Terminology
Abstract
Mechanistic interpretability identifies sparse subsets of heads and MLP blocks that carry specific behaviors. We ask whether such causal signals can guide where to place a small PEFT budget more effectively than the cheap heuristics practitioners already use. scores attention heads and MLP blocks by normalized activation-patching recovery on clean/corrupted probes and trains LoRA/QLoRA only on the selected sites; adds bounded credit for small within-layer joint subsets. We compare against random, magnitude, activation-norm, and gradient/Fisher on Ministral-8B/NF4 in three cells: Swahili span-JSON information extraction (IE) at b = 0.25% and 1.0%, and English to Swahili machine translation (MT) at b = 1.0%. The causal selectors never win the primary metric. On the headline IE cell (3 seeds, paired-bootstrap CIs over 600 predictions), beats random by +0.079 span+type F1 and gradient/Fisher by +0.174, but trails activation-norm by 0.028, with the smallest cross-seed std (plus or minus 0.003). On MT all four selectors lie within 0.30 BLEU and every paired CI includes zero. A schema-versus-span decomposition explains the IE gap: activation-norm captures the rigid JSON routine, while causal scores track content-sensitive sites. We distill a preliminary diagnostic -- prefer activation-norm when output structure dominates, treat causal selectors as a hypothesis for content-dominated tasks -- and release masks, scores, predictions, and evaluation files for direct replay.
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering