Perturbation-based Regional Interpretability through Subtraction Mapping (PRISM): naming-error dissociations in language models and post-stroke aphasia
Xiang Guan, Roger D. Newman-Norlund, Yong Yang, Saeed Ahmadi, Regan Willis, Nadra Salman, Kalil Warren, Srihari Nelakuditi, Chris Rorden, Leonardo Bonilha, Julius Fridriksson
University of South Carolina · ALLT.AI, LLC · USC School of Medicine
cs.LG, cs.CL
Submitted: 2026-08-13
Updated: 2026-08-14
Comments: 49 pages, 6 figures, 1 table. Supplementary methods, 6 tables and 5 figures included
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
Importance score: 75/100
The gist: PRISM (Perturbation-based Regional Interpretability through Subtraction Mapping) adapts subtraction analysis, the standard framework of human neuroimaging, from biological brains to perturbed
Terminology
Summary
PRISM (Perturbation-based Regional Interpretability through Subtraction Mapping) adapts subtraction analysis, the standard framework of human neuroimaging, from biological brains to perturbed transformers, and applies the same logic to both substrates in parallel. Building on the Brain-LLM Unified Model (BLUM), which showed that layer-perturbed LLaVA-1.6-Vicuna-13B error profiles match the lesion patterns of aphasic patients, PRISM maps the seven clinical Philadelphia Naming Test categories, subtracts error classes pairwise, and treats each perturbation seed as a subject in a group analysis with threshold-free cluster enhancement along the layer axis. A structurally matched analysis is run on 213 chronic post-stroke aphasia patients using correlation-difference lesion-symptom mapping, and both sides are replicated on held-out splits. The designs match in subject dimension (seeds, patients), spatial dimension (layers, atlas-parcellated cortex) and thresholding, but the contrast operator differs: a within-subject error-proportion difference for the LLM, a between-subject correlation difference for the cortex. Both substrates recover a robust phonemic-favoring dissociation, a deep layer cluster and a frontal-perisylvian cortical cluster, both replicating; the semantic-favoring direction is a consistently signed but non-significant trend on both. PRISM thus gives a falsifiable, spatially resolved test of functional-specialization claims in transformer language models. A confirmatory ROI-level intervention (PRISM Stage 3) licensing the strongest causal-mechanism claim is left to subsequent work.
On the LLM side, the primary Semantic-versus-Phonemic contrast under seed-as-subject layer-axis TFCE yields a robust Phonemic > Semantic layer cluster at layers 22–31 (discovery, validation 23–33), which is corroborated by a coincident Phonemic > Neologism cluster at layers 24–31 against the non-real-word baseline. The Semantic > Phonemic direction does not survive on the LLM in any of the three confirmatory contrasts. PRISM Stage 2 layer-order permutation testing places every observed cluster at empirical p < 0.001 against a null in which layer order is destroyed within each seed. On the cortex, the matched ROI-level correlation-difference VLSM with 50/50 patient split recovers the same asymmetry: PoCG L and PrCG L survive the strictest replication criterion in the Phonemic > Semantic direction on both halves; the Semantic > Phonemic direction has consistently signed point estimates that do not survive replication. The honest cross-substrate headline is therefore that the phonemic-favoring direction is robust and replicates on both substrates, and the semantic-favoring direction is a consistently signed but non-significant trend on both substrates, an asymmetry that is itself a replicable property of the data rather than a failure of the method.
Each surviving phonemic-favoring layer cluster on the LLM side carries a dose-response signature with respect to perturbation depth (σ) and perturbation extent (ρ), the LLM analogs of lesion severity and lesion extent. For the primary Semantic-vs-Phonemic cluster (layers 22–31), the discovery-seed OLS dose-response slopes are βσ = −0.182 (95% cluster-bootstrap CI [−0.185, −0.179]) and βρ = −0.462 ([−0.474, −0.451]), both with the same sign as the cluster’s contrast: deeper and more extensive perturbations produce a stronger Phonemic > Semantic dissociation on average. The same slopes on validation seeds are essentially identical (βσ = −0.176, βρ = −0.434). The Phonemic-vs-Neologism cross-check shows the same scaling pattern at coincident layers (24–31); the Semantic-vs-Neologism Neologism > Semantic effect at the early mid-network layers (19–20) scales in the same (negative) direction as its contrast, so the Neologism > Semantic dissociation likewise strengthens with both σ and ρ.
The dose-response surface is informative beyond the linear slope. Across all three clusters, the per-cell mean D does not increase monotonically with σ and ρ across the full grid. It peaks along a diagonal ridge on which perturbation depth and perturbation extent trade off inversely, and falls away on both sides of that ridge: for the primary cluster the global maximum sits at σ = 1.5, ρ = 0.7 (D = 0.22), with comparable maxima at σ = 1.8, ρ = 0.4 and σ = 2.0, ρ = 0.3, while both the low-dose corner (σ = 1.1, ρ = 0.1) and the high-dose corner (σ = 2.0, ρ = 1.0) fall to near zero. What governs the contrast is therefore the total amount of perturbation delivered rather than either parameter alone. This is interpreted as a category-differentiation regime: moderate total perturbation maximally separates error categories from each other, while severe perturbation degrades that differentiation as all error categories rise together. The directional slopes (βσ and βρ, each matching its cluster’s contrast sign) are coarse directional summaries of the average dose dependence across the grid; the bilinear-plus-interaction OLS explains only a modest share of each surface (R2 ≈ 0.18–0.26), so the slopes are treated as directional descriptors of the observed σ × ρ surfaces rather than as a global model of them. The saturating shape captures the limit of the perturbation-as-lesion analogy at very high perturbation levels. Both properties replicate on the held-out 40-seed validation cohort.
PRISM demonstrates that the subtractive approach can be applied both to a 13-billion-parameter vision-language transformer and to a cohort of 213 chronic-aphasia patients administered the same task, with the parallel holding at the level of subject (seeds for the LLM, patients for the cortex), spatial dimension (transformer layers for the LLM, atlas-parcellated cortex for the patients), and TFCE-style thresholding along an ordered axis. The contrast operator is not parallel, the LLM side is a within-subject difference in error proportions averaged across seeds, the human side is a between-subject correlation difference, but the spatial-inferential machinery is. The Semantic-versus-Phonemic contrast and its two Neologism cross-checks each recover a robust phonemic-favoring map on both substrates, a frontal-perisylvian Phonemic > Semantic ROI cluster on the cortex (PoCG, PrCG, SLF; replicating on a 50/50 patient split) and a deep Phonemic > Semantic layer cluster on the LLM (layers 22–31; replicating on a 40/40 seed split and clearing Stage-2 layer-permutation at p < 0.001). The semantic-favoring direction is a consistently signed but non-significant trend on both substrates, an asymmetry that itself replicates and is the honest characterization of the data. Demonstrating both substrates within a single study, with each analyzed at the inference unit appropriate to its data, establishes the methodological prerequisite for patient-specific computational surrogates of neurological language disorders, with applications to clinical-trial design, individualized rehabilitation, and pharmacological prediction.
Improvements for AI systems
Improvements to AI systems:
-
Perturbation-based interpretability with dose-response modeling: Implement PRISM’s subtractive error-class analysis to map functional specialization in transformer layers. The system can now quantify how perturbation depth (σ) and extent (ρ) affect task-specific error dissociations (e.g., phonemic vs. semantic), revealing non-monotonic, ridge-shaped dose-response surfaces that identify optimal perturbation regimes for category differentiation—useful for diagnosing model failure modes and calibrating intervention strength.
-
Seed-as-subject group statistics with TFCE thresholding: Adopt seed-as-subject group analysis (treating each random seed as an independent subject) with threshold-free cluster enhancement along the layer axis. This enables robust, replicable identification of functionally specialized layer clusters (e.g., layers 22–31 for phonemic processing) with empirical p-values from layer-order permutation tests, improving confidence in causal claims about where and how knowledge is organized in LLMs.
-
Cross-substrate parallel validation framework: Apply the same spatial-inferential machinery (subject dimension, ordered spatial axis, TFCE) to both LLMs and human neuroimaging data (e.g., lesion-symptom mapping). The improved system can now validate AI internal representations against biological brain lesion patterns, enabling falsifiable tests of functional-specialization claims and identifying which model layers correspond to specific cognitive deficits (e.g., phonemic vs. semantic aphasia).
-
Dose-response surface modeling for perturbation severity: Use the bilinear-plus-interaction OLS with cluster-bootstrap confidence intervals to characterize how perturbation parameters trade off (e.g., depth vs. extent) in producing maximal category dissociation. The system can now predict the total perturbation dose required to elicit a desired error profile, and detect saturation where further perturbation degrades differentiation—useful for adversarial robustness testing and fine-tuning strategies.
-
Replicable asymmetry detection: Implement the finding that semantic-favoring directions are consistently signed but non-significant, while phonemic-favoring directions are robust and replicable. The improved system can now flag one-sided functional asymmetries in model representations, preventing over-interpretation of weak trends and guiding targeted interventions only where evidence is strong.
What the improved AI system can do:
-
Localize functional modules: Automatically identify which transformer layers encode phonemic vs. semantic vs. neologism distinctions, with statistical rigor (cluster-level p < 0.001) and cross-seed replication.
-
Predict lesion-like effects: Given a perturbation (e.g., pruning, noise, or targeted layer damage), predict the resulting error category dissociation and its magnitude, using dose-response surfaces calibrated on seed cohorts.
-
Bridge AI and neuroscience: Compare model layer clusters to cortical ROI clusters (e.g., PoCG, PrCG, SLF) in aphasia patients, enabling patient-specific computational surrogates for language disorders—useful for simulating treatment outcomes or pharmacological effects before clinical trials.
-
Optimize perturbation protocols: Determine the minimal perturbation dose (σ, ρ) that maximizes category differentiation without saturating, improving interpretability tools like saliency maps or counterfactual explanations.
-
Validate causal claims: Use layer-order permutation testing and held-out seed splits to confirm that identified clusters are not artifacts of seed variance or layer ordering, ensuring that mechanistic claims about model internals are statistically defensible.
Abstract
Mechanistic interpretability of large language models lacks spatially resolved, falsifiable tools for testing whether internal components are specialized for distinct cognitive operations. We adapt subtraction analysis, the standard framework of human neuroimaging, from biological brains to perturbed transformers, and apply the same logic to both substrates in parallel. Building on the Brain-LLM Unified Model (BLUM), which showed that layer-perturbed LLaVA-1.6-Vicuna-13B error profiles match the lesion patterns of aphasic patients, we develop PRISM (Perturbation-based Regional Interpretability through Subtraction Mapping). PRISM maps the seven clinical Philadelphia Naming Test categories, subtracts error classes pairwise, and treats each perturbation seed as a subject in a group analysis with threshold-free cluster enhancement along the layer axis. We run a structurally matched analysis on 213 chronic post-stroke aphasia patients using correlation-difference lesion-symptom mapping, and replicate both sides on held-out splits. The designs match in subject dimension (seeds, patients), spatial dimension (layers, atlas-parcellated cortex) and thresholding, but the contrast operator differs: a within-subject error-proportion difference for the LLM, a between-subject correlation difference for the cortex. Both substrates recover a robust phonemic-favoring dissociation, a deep layer cluster and a frontal-perisylvian cortical cluster, both replicating; the semantic-favoring direction is a consistently signed but non-significant trend on both. PRISM thus gives a falsifiable, spatially resolved test of functional-specialization claims in transformer language models. A confirmatory ROI-level intervention (PRISM Stage 3) licensing the strongest causal-mechanism claim is left to subsequent work.
Sources
- Generative causal testing to bridge data-driven models and scientific theories in language neuroscience
- Sparse Autoencoders Find Highly Interpretable Features in Language Models
- Stroke Lesions as a Rosetta Stone for Language Model Interpretability
- How to use and interpret activation patching
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks