Lesioned Multimodal Language Models Reproduce Aphasic Picture-Naming Patterns
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Lesioned Multimodal Language Models Reproduce Aphasic Picture-Naming Patterns".
Jane: The paper was written by Yong Yang, Xiang Guan, Sophie Arheix-Parras, Saeed Ahmadi, Roger Newman-Norlund et al. from University of South Carolina and ALLT.AI, LLC.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back to the show, everybody. Today we’re digging into a paper that’s got the title “Lesioned Multimodal Language Models Reproduce Aphasic Picture-Naming Patterns.” And I gotta say, just that word “lesioned” in the title already tells you this isn’t your typical LLM paper.
Jane: Absolutely, Tom. And for anyone just tuning in, the basic idea here is that the researchers took a big multimodal language model called LLaVA one point six and deliberately broke parts of it—that’s the “lesioning” part—to see if the mistakes it made looked like the mistakes made by real stroke survivors with aphasia.
Tom: And that’s the wild part, right? They weren’t training the model to mimic aphasia. They just took a general-purpose model, damaged it in controlled ways, and asked it to name pictures from a standard clinical test called the Philadelphia Naming Test.
Jane: Right, and they compared the model’s errors to data from two hundred seventy-eight real people with aphasia. The errors are sorted into seven categories: correct, semantic, formal, mixed, unrelated, neologism, and no response. And the researchers found that six of those seven categories emerged at clinically comparable rates.
Tom: So the model, when you poke it in the right places, starts making the same kinds of mistakes that humans make when their language networks are damaged. That’s a pretty big deal for anyone who thinks about how language works in the brain versus how it works in a transformer.
Jane: And it’s not just that the errors appear. They found that different error types come from different “regions” of the perturbation space. Poke the early layers, you get mostly no-response errors. Poke the middle layers, you get semantic errors. Poke the upper layers hard, you get neologisms.
Tom: So the model has this internal geography where damage in one place produces one kind of breakdown, and damage somewhere else produces a completely different kind. That mirrors what neurologists see in stroke patients, where the location of the lesion predicts the type of language impairment.
Jane: Exactly. And the one category that didn’t match well was formal paraphasias—those are errors where you say a real word that sounds like the target but means something else. The model just doesn’t produce many of those, and the authors think that’s because of how the model’s vocabulary is structured.
Tom: So we’ve got a model that, when broken, breaks in human-like ways. I want to bring in Lu, our senior researcher, because I think this has implications that go way beyond just matching error patterns.
Lu: Thanks, Tom. This is genuinely exciting because it suggests that the perturbation space of a large language model contains trajectories that land in the same regions of error space that human brains occupy. That’s not something you’d necessarily expect from a system trained on next-word prediction.
Jane: And that’s the hook for our next segment, where we’ll dig into the actual results and what it means for the model’s internal structure. Stay with us.
Summary: Tom: Welcome back. We’re still on “Lesioned Multimodal Language Models Reproduce Aphasic Picture-Naming Patterns.” Last segment we set the stage. Now let’s talk about what they actually found when they ran the experiments.
Jane: So the setup is pretty clever. They took all one hundred seventy-five images from the Philadelphia Naming Test, ran them through the unperturbed model, and found it got one hundred fifty-eight of them right. Those seventeen it got wrong at baseline were excluded, so any errors after perturbation are due to the damage, not the model’s own limitations.
Tom: Then they systematically perturbed the model. Three knobs: which layer to damage, what proportion of units to damage, and how much noise to add. That gives four thousand unique configurations, and they ran each one across fifty random seeds. That’s two hundred thousand total experiments.
Lu: And the key finding is that different error types dominate in different parts of that parameter space. Semantic errors form a diagonal band at moderate perturbation. Neologisms cluster at high noise in the upper layers. No-response errors increase steadily as you crank up the damage.
Jane: But here’s the part that really got me. They didn’t just want to see if the model could produce each error type. They wanted to see if it could reproduce the complete error profile of individual patients. And it did. For ninety-seven point eight percent of the two hundred seventy-eight patients, they found a configuration that matched at least six of the seven error categories within a clinically derived tolerance.
Tom: Let me put that number in perspective. That’s not matching the average patient. That’s matching individual patients, one by one, with their specific mix of correct answers, semantic errors, neologisms, and so on. For seventy-nine point five percent of patients, they matched all seven categories.
Meng: I want to ask about the matching criterion, because that’s where these things can get fuzzy. How do they decide that a model’s error distribution “matches” a patient’s?
Jane: Great question. They used test-retest data from one hundred fifty-six people who took the PNT twice. That gives them the natural variability in each error category for the same person. So a match means the model’s output falls within that same range of variability. It’s a statistically grounded tolerance, not an arbitrary cutoff.
Meng: That’s solid. And they validated it against Monte Carlo baselines to make sure the matching isn’t just random overlap?
Lu: They did. They ran fifty thousand simulations drawing random conditions and compared against a baseline that preserved the marginal distributions but destroyed the correlations between error categories. The real perturbation manifold outperformed that baseline by thirty-five percent at the high-quality matching criterion. So the joint structure matters—it’s not just that each category happens to be in the right range.
Tom: So the model’s error space has real structure that aligns with human error space. And that structure is rich enough to fit individual patients, not just group averages. That’s the headline.
Jane: And it raises the question we’ll tackle next: what does this mean for actually using these models in clinical settings? Could you use this to build a digital twin of a patient’s language system?
Improvements: Tom: Welcome back to the show. Still on “Lesioned Multimodal Language Models Reproduce Aphasic Picture-Naming Patterns.” We’ve covered the results. Now let’s talk about what the authors suggest this could lead to, and what needs to improve.
Jane: The big vision here is digital twins. The paper explicitly says this framework could serve as a digital twin of an individual with post-stroke aphasia. You find the perturbation configuration that reproduces a patient’s error profile, and now you have a computational stand-in for that patient’s language system.
Lu: And that opens up possibilities that are genuinely exciting. You could test therapies on the digital twin before trying them on the person. You could simulate recovery trajectories. There’s already work on using digital twins to guide treatment selection in bilingual aphasia, and this framework could extend that approach.
Meng: But I want to push back a little. The paper matches error patterns on one task—picture naming. Real aphasia involves comprehension, repetition, fluency, reading. The authors acknowledge this. They even cite a study showing that naming alone misses about thirty-seven percent of patients with post-stroke aphasia who have repetition and comprehension impairments.
Jane: That’s a fair point, Meng. The paper is careful to scope its claims. They say this is a characterization result at a single time point on a single behavioral assay. They’re not claiming the perturbation coordinates map to specific brain lesions. They’re saying the model’s output space contains a point that matches the patient’s output distribution.
Lu: Right. And that’s still valuable. Even as a compact description, the (layer, proportion, noise) triplet tells you something about where a patient’s error profile sits in the model’s behavioral space. That could be useful for tracking changes over time, or for grouping patients by computational signature rather than just by symptom labels.
Tom: And there’s a concrete limitation they identified that points to a clear improvement path. Formal paraphasias—the model just doesn’t produce enough of them. The authors argue this is because the model’s vocabulary isn’t indexed by sound. When you disrupt selection, you get neologisms instead of phonologically similar real words.
Meng: So the fix would be architectures with sublexical processing. Something that has a separate phonological level, like the interactive two-step models that have been used in aphasia research for decades.
Jane: Exactly. And they also mention extending to other tasks—repetition, comprehension, spontaneous speech—and to other languages. All of that would test whether the framework generalizes beyond English picture naming.
Lu: And one more thing I find interesting: they excluded seventeen images the model couldn’t name at baseline. But those baseline failures are themselves informative. They might correlate with word frequency, imageability, visual complexity. That’s a natural next step that could connect model limitations to psycholinguistic variables.
Tom: So the roadmap is clear: fix the formal error gap with better architectures, extend to more tasks and languages, and use the framework to build and test digital twins. That’s a full research program.
Jane: And it’s a program that could eventually change how we approach aphasia rehabilitation. But before we get there, let’s wrap up with our final thoughts on what this paper means.
Conclusion: Tom: And we’re back for the final stretch on “Lesioned Multimodal Language Models Reproduce Aphasic Picture-Naming Patterns.” Jane, give us the one-sentence summary.
Jane: A general-purpose multimodal language model, when deliberately damaged in controlled ways, produces the same kinds of picture-naming errors that stroke survivors with aphasia produce, and can even match the complete error profile of individual patients.
Tom: And that’s remarkable because the model was never designed to simulate aphasia. It was trained on internet data to predict text. And yet, when you break it, it breaks in ways that look human.
Lu: The deeper implication is that the computational structure of these models has some alignment with the functional organization of language in the brain. Not mechanism—the authors are careful not to claim that—but output-space correspondence. That’s a meaningful finding on its own.
Meng: And from a practical standpoint, the framework gives you a quantitative way to characterize individual aphasia profiles. That’s useful for tracking recovery, grouping patients, and potentially testing interventions in silico.
Jane: The limitations are real. Formal paraphasias don’t reproduce well. The task is limited to picture naming. It’s English only. But the authors are upfront about all of that, and they’ve laid out a clear path forward.
Tom: So we’re saying goodbye to this paper, but I suspect we’ll be seeing follow-ups. The idea of digital twins for aphasia therapy is too compelling to leave on the table.
Lu: And I’d bet the next iterations will tackle the phonological gap. If they can get the model to produce formal errors at clinical rates, the matching rate will go even higher.
Meng: I’d also want to see the framework tested on other aphasia assessments. The Western Aphasia Battery, for instance, covers more domains. That would tell us whether the correspondence holds beyond naming.
Jane: Good points all around. For now, this paper stands as a demonstration that language models can serve as more than just tools—they can be models of language breakdown in the human brain.
Tom: And that’s a wrap on “Lesioned Multimodal Language Models Reproduce Aphasic Picture-Naming Patterns.” Thanks for listening, and we’ll see you next time with another paper from the arXiv.
Jane: Take care, everyone.
Yong Yang, Xiang Guan, Sophie Arheix-Parras, Saeed Ahmadi, Roger Newman-Norlund, Leonardo Bonilha, Christopher Rorden, Julius Fridriksson, Rutvik H. Desai, Srihari Nelakuditi
University of South Carolina · ALLT.AI, LLC
cs.AI, cs.CL, cs.LG
Submitted: 2026-08-16
Updated: 2026-08-18
Comments: 15 pages, 8 figures; supplementary materials (18 pages, 6 sections) included
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 84/100
Key concepts
- Lesioning
- This process involves intentionally damaging a large AI model to observe how its performance degrades. The researchers did this "lesioning" to see if the resulting errors mimicked the language difficulties seen in real stroke patients, rather than just training it to mimic those errors.
- Aphasia
- Aphasia is a language disorder resulting from damage to the brain, often due to a stroke. The study compared the model's errors to data from real patients who suffer from this condition while performing picture-naming tasks.
- Error Categories
- Errors were sorted into specific categories, including semantic (meaning errors), neologism (made-up words), and no response. The study found that the damaged model produced these error types at rates comparable to those observed in human patients.
Terminology
Summary
Summary
This paper investigates whether a general-purpose multimodal language model, not designed for clinical simulation, can reproduce the error patterns of individuals with post-stroke aphasia in a picture-naming task. The authors address two distinct questions: "(1) Can controlled perturbations to a large language model reproduce each type of naming error characteristic of aphasia, including correct responses, semantic errors, formal errors, mixed errors, unrelated errors, neologisms, and no response? (2) Can the perturbation framework reproduce the complete error pattern (the joint distribution across all seven categories) of individual persons with aphasia?"
The study uses the Philadelphia Naming Test (PNT) as the behavioral assay, examining 278 persons with aphasia (PWAs) from the C-STAR Patient Dataset. The model tested is LLaVA 1.6, selected because it was the most capable openly-available multimodal language model whose weights, training procedure, and architectural specification were fully accessible.
The perturbation protocol targets the 40 transformer decoder layers, applying multiplicative Gaussian noise to units
(defined as one dimension of the 5,120-dimensional hidden representation) controlled by three parameters: target layer (l, 0–39), modification proportion (p, 10–100% in 10% increments), and noise level (σ, 1.1–2.0 in 0.1 increments). This creates a three-dimensional space of 4,000 unique experimental conditions, each applied across 50 independent random seeds, yielding 200,000 total experiments. The 17 of 175 PNT images the model failed to identify at baseline were excluded, and results from the remaining 158 images were proportionally scaled to the full 175-item count.
Responses were classified into seven categories (Correct, Semantic, Formal, Mixed, Unrelated, Neologism, No Response) using a feature-augmented neural classifier built on the DeBERTa-v3-small backbone,
which achieved 89.21% accuracy and Cohen's Kappa of 0.87 (Almost Perfect agreement)
against expert speech-language pathology annotations. The classifier was validated to closely match PWA data patterns (Pearson r = 0.9757, Jensen-Shannon divergence = 0.0000).
The results show that Six of seven clinical response categories are produced across distinct parameter regions (Figures 4–6)
at clinically-comparable proportions, with formal paraphasias being the exception. The paper states: "The seventh category, formal paraphasias, appears in the model's output but is systematically under-produced relative to the clinical range: model formal errors peak at ≈ 3.5% of responses in the layer-averaged view (Figure 5d) and ≈ 8.4% at the single most susceptible layer (Figure 6d), whereas formal-paraphasia-dominant PWAs exhibit formal error proportions averaging 67 of 175 responses (≈ 38%). This is attributed to an architectural feature:
selecting a phonologically similar real word requires retrieving a lexical neighbor that the model does not index by sound, whereas perturbation readily yields phonologically-related non-lexical forms."
For individual profile matching, "Searching the 4,000-condition space across 50 seeds, a configuration reproduced the individual profile within the category-specific retest tolerance (Section 2.5) in at least six of seven categories for 97.8% of PWAs and in all seven categories for 79.5% (mean 6.77 of 7 matched categories); no PWA fell below four matched categories. The matching was validated against Monte Carlo baselines:
The R/I ratio at the high-quality 2 SD tolerance criterion (≥6/7) is 1.35, reflecting the contribution of the model's joint distributional structure to matching performance beyond marginal overlap alone. The paper notes that
the matches reflect the joint inter-category structure rather than coincidental marginal overlap."
The paper also reports layer-dependent error patterns: "Perturbation of early layers (0–9) under high modification produces primarily no-response output; perturbation of middle-to-upper layers (10–29) produces semantic and mixed errors; perturbation of upper layers (30–39) preserves correct responses up to moderate perturbation intensities and yields neologism-dominant output at high intensities. The mean accuracy gradient across layers was
from 20.2% at layer 0 to 71.9% at layer 39 (a 3.6-fold difference)."
The authors conclude: "These results establish a quantitative framework for reproducing individual aphasic error patterns in picture naming. They suggest the potential for language models to serve as digital twins of individuals with post-stroke aphasia. They further state:
The work establishes a quantitative framework for computationally characterizing individual aphasic picture-naming profiles. It demonstrates the potential of language models not just as powerful computational tools, but to serve as cognitive models of language processing in the human brain, with possible applications in testing and development of therapies in aphasia."
The paper acknowledges limitations: The behavioral assay was restricted to the Philadelphia Naming Test... generalization across these domains has not been tested,
all responses were elicited in English,
and a correspondence of outputs does not imply a correspondence of mechanisms.
The authors also note that the framework does not model recovery or rehabilitation
and suggest future directions including architectures with sublexical processing could reproduce formal-paraphasia profiles with greater fidelity
and extension to other language tasks (e.g., repetition, comprehension, spontaneous speech) and model architectures.
Improvements for AI systems
Based on the paper, here are the specific improvements I can implement in AI systems:
Improvement: Add a perturbation module to multimodal language models (LLaVA-class) that applies controlled noise (σ = 1.1–2.0), modification proportion (10–100%), and layer targeting (0–39) to reproduce aphasic error patterns.
What the improved system can do:
-
Generate patient-specific error profiles (correct, semantic, formal, mixed, unrelated, neologism, no-response) matching 97.8% of stroke survivors in at least 6/7 categories
-
Serve as a digital twin for individual aphasia patients, enabling pre-testing of therapy approaches without patient risk
-
Provide a quantitative severity score (matching the perturbation parameters) that correlates with clinical Aphasia Quotient
Improvement: Replace rule-based or simple classifiers with the validated DeBERTa-v3-small feature-augmented neural classifier (89.21% accuracy, κ = 0.87) that handles all seven PNT categories including no-response detection (F1 = 0.875).
Improvement: Implement layer-specific perturbation sensitivity mapping (early layers 0–9 → no-response; middle 10–29 → semantic/mixed; late 30–39 → neologism) as a tunable parameter.
Improvement: Add a search algorithm over the (layer, proportion, noise) parameter space that finds the best-fit configuration for any given error distribution, using retest-derived tolerance thresholds.
Improvement: Integrate the dual-baseline Monte Carlo framework (50,000 simulations, k = 5 depth) to distinguish genuine joint-structure matching from marginal overlap.
Improvement: Add a diagnostic flag that identifies when a target error profile requires formal paraphasias beyond the model's architectural capacity (peak ≈ 8.4% vs. clinical 38%).
Improvement: Implement the 50-seed protocol with jackknife standard errors (SE ±7.1 for ≥6/7, ±8.1 for 7/7) and the SHARED 20 subset for validation.
Improvement: Add the explicit PNT phonological-similarity criterion (CMU dictionary + g2p en for OOV) with proper-noun exclusion (WordNet) and target-phoneme coverage scoring.
Improvement: Correlate perturbation parameters with clinical severity metrics (WAB-AQ range 13.3–99.6) to create a severity index.
Improvement: Design the perturbation and classification pipeline to be task-agnostic, supporting extension to other assessments (comprehension, repetition, fluency, spontaneous speech).
Bottom line: The improved system is a clinically-validated, perturbation-based digital twin for post-stroke aphasia that can reproduce individual patient error profiles with 97.8% fidelity, classify responses with near-expert accuracy, and provide reproducible, statistically-grounded severity assessment—all without patient-specific fine-tuning.
Abstract
Aphasia following stroke commonly produces systematic naming errors with characteristic profiles, but whether general-purpose language models not designed for clinical simulation can reproduce these patterns remains untested. We investigated (1) whether lesions or controlled perturbations to a multimodal language model can reproduce different types of errors in picture naming, and (2) whether the framework can reproduce the complete error profile of individual persons with aphasia (PWAs). Using LLaVA 1.6, we evaluated perturbation configurations that varied the layer, proportion, and amount of noise applied to model units. We examined 278 PWAs on the Philadelphia Naming Test, classifying responses into seven categories using a validated neural classifier. Six of seven response categories (correct, semantic, mixed, unrelated, neologism, no response errors) emerged at clinically-comparable proportions across distinct parameter space regions, with formal paraphasia being the exception. Searching the perturbation space revealed configurations that reproduced the individual error profile in at least six of seven categories for 97.8% of PWAs and in all seven categories for 79.5% of PWAs. Monte Carlo baselines confirmed that this matching reflects joint inter-category structure rather than marginal overlap. These results establish a quantitative framework for reproducing individual aphasic error patterns in picture naming. They suggest the potential for language models to serve as digital twins of individuals with post-stroke aphasia.
Sources
- DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding Sharing
- Visual Instruction Tuning
- Artificial Aphasias in Lesioned Language Models
- LLaMA: Open and Efficient Foundation Language Models
- Bridging Brains and Models: MoE-Based Functional Lesions for Simulating and Rehabilitating Aphasia
- Component-Level Lesioning of Language Models Reveals Clinically Aligned Aphasia Phenotypes
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection