From Observation to Intervention: Memory in Brains and Large Language Models
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "From Observation to Intervention: Memory in Brains and Large Language Models".
Jane: The paper was written by Morteza Salehjahromi, Shayan A. Zadegan, Amgad Muneer and Jia Wu from The University of Texas MD Anderson Cancer Center and The University of Tennessee Health Science Center.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back to the channel, everyone. Today we're looking at a paper that's been making the rounds on arXiv, and it's called "From Observation to Intervention: Memory in Brains and Large Language Models." Jane, I gotta say, the title alone got me hooked.
Jane: Oh, same here, Tom. It's one of those titles that tells you exactly what the authors are trying to do. They're not saying brains and AI models are the same thing. They're saying we can ask the same kinds of questions about both, and that's a much more useful framing.
Tom: Right, because the moment you say "artificial hippocampus" or something like that, you're already in trouble. The authors actually call that out on the very first page. They say a model component shouldn't be called an artificial hippocampus just because it participates in retrieval.
Jane: Exactly. And that's the core of the paper. The authors are from MD Anderson Cancer Center and the University of Tennessee, and they're basically arguing that we've spent decades watching brains form memories, but we can't easily poke and prod the specific neurons involved. LLMs, on the other hand, we can edit their weights, change their activations, swap out their memory stores, and rerun the same experiment a thousand times.
Tom: So the brain gives us rich observation, but limited intervention. The LLM gives us the opposite.
Jane: You got it. And the authors think that's actually a huge opportunity. Instead of always going from biology to AI, they want to reverse the pipeline. Start with an intervention in an LLM, see what happens, and then design a sharper biological experiment based on that.
Tom: That's a bold idea. I mean, we've had neuroscience inspiring AI for decades. Now they're saying AI can inspire neuroscience.
Jane: And they're careful to say it's not because LLMs have better memory. They don't have lived experiences. They don't have emotions or a body. But they do have something no brain researcher has ever had: total control over the system.
Tom: So the title is really about moving from just watching to actually changing things, and doing that in both fields.
Jane: Precisely. And the authors lay out four shared questions that structure the whole paper. Representation, retrieval from partial cues, writing and updating, and intervention. We're going to dig into each of those as we go through the pages.
Tom: I'm already excited for that. And I think our listeners are going to love the part where they compare what we can actually do in humans, rodents, and LLMs. That table alone is worth the read.
Jane: Oh, definitely. But before we get there, let's talk about what the paper actually summarizes from the neuroscience side. That's where the story really starts.
Summary: Tom: So, Jane, we're back with "From Observation to Intervention: Memory in Brains and Large Language Models." Let's talk about what the paper actually summarizes from decades of memory research.
Jane: It's a beautiful summary, honestly. The authors walk us through the history. You start with hippocampal indexing theory from the 1980s, where the hippocampus links information scattered across the cortex. Then you get complementary learning systems, which explains why we can learn a new fact instantly but still integrate it slowly into general knowledge.
Tom: And then they bring in the single-neuron recordings. That's the stuff that really blew my mind when I first read about it.
Jane: Oh, absolutely. There's this work by Quiroga and Fried showing neurons that respond to specific concepts, like Jennifer Aniston, across completely different images. Then Gelbard-Sagiv showed those same neurons reactivating before a person even says what they remember. And Ison showed neurons rapidly forming new person-place associations after just a few exposures.
Tom: So the brain is doing all this incredible coding at the single-cell level. But here's the catch. You can watch it, but you can't easily poke it.
Jane: Right. And that's where the intervention gap comes in. In rodents, you can tag engrams and optogenetically activate or silence them. That's the Liu and Tonegawa work. But in humans and macaques, interventions are usually much broader. You're stimulating the whole entorhinal region or the whole hippocampus, and the effects depend on timing, frequency, and the task.
Tom: And sometimes stimulation helps memory, sometimes it hurts it. The Suthana study showed enhancement, the Jacobs study showed impairment.
Jane: Exactly. So the summary is that we have this rich phenomenology from human recordings, but we're stuck at the observational level. Rodents give us more causal access, but less behavioral complexity. And then LLMs come along and just flip the whole thing.
Tom: Because in an LLM, you can locate a specific factual association, edit it, and measure the side effects. The paper cites ROME, which is a method for editing facts in GPT. You can steer activations with ActAdd. You can shift representation directions with representation engineering.
Jane: And you can do all of that while keeping everything else fixed. Same model, same prompt, same conditions. Just one change at a time. That level of control is unprecedented.
Tom: So the summary is really about this asymmetry. Brains give us observation. LLMs give us intervention.
Jane: And the authors argue that this asymmetry is the opportunity. Not because LLMs are better at memory, but because they're better at letting us experiment. And that's what the rest of the paper builds on.
Tom: So what do they actually propose we do with that advantage?
Jane: That's the next segment. They've got a whole framework for reversing the experimental pipeline.
Improvements: Tom: We're continuing with "From Observation to Intervention: Memory in Brains and Large Language Models." And Jane, I think this is where the paper really gets exciting. They're not just describing the gap. They're proposing a way to use it.
Jane: Right. The improvement they suggest is essentially a reverse experimental pipeline. Instead of always going from human observation to animal studies to AI models, they want to start with LLM interventions and then test those findings in biological systems.
Tom: So you find something interesting in a model, and then you go look for it in a mouse.
Jane: Exactly. And they give a concrete example. There's this method called causal tracing, from the ROME paper, where you identify which layer in a transformer is actually responsible for a factual association. You corrupt the input, then restore activations at different layers to see where the fact lives.
Tom: And they compare that to human recordings showing neurons reactivating before recall. But the full sequence from cue to answer is unclear in both systems.
Jane: Right. So the authors propose that if an outdated association appears early in the LLM and remains detectable even after the updated answer becomes dominant, you could test whether the brain similarly reactivates an older memory before it gets suppressed.
Tom: That's a testable hypothesis. You could look for that in human intracranial recordings.
Jane: And they also propose a shared evaluation language for memory updating. LLM editing research already checks whether an edit generalizes, whether it preserves unrelated knowledge, whether it lasts over time, whether it causes unintended effects. The authors say biological memory studies should use the same criteria.
Tom: So if you're changing a context from threat to safety in a rodent, you should ask whether that update generalizes to similar cues, whether it leaves unrelated threat memories alone, whether it persists, whether it comes back under stress.
Jane: Exactly. And they're careful to say this doesn't mean the mechanisms are the same. It just means we can evaluate interventions in a comparable way.
Tom: I love that. It's like a benchmark for memory editing across species and systems.
Jane: And they add guardrails. An LLM result is a hypothesis generator, not evidence about the brain. You shouldn't map brain regions to model components. You should test for unintended effects. And null results are informative.
Tom: So if a principle works in an LLM but fails in a mouse, that tells you something about biological memory.
Jane: It tells you that the biological system has constraints the model doesn't have. The body, emotion, neuromodulation, lived experience. All of that shapes memory in ways a transformer will never capture.
Tom: So the improvement is really about building a two-way street. Neuroscience inspires AI, and now AI can inspire neuroscience.
Jane: And they think this could lead to discoveries that aren't obvious from observation alone. That's the promise.
Tom: I want to get into the actual first page now, because there's a lot of nuance there about how they define terms.
Jane: Good call. Let's look at that.
First Page: Tom: Alright, we're back with "From Observation to Intervention: Memory in Brains and Large Language Models." Let's actually dig into the first page, because there's a lot packed in there.
Jane: The first page sets up the whole framing. The authors start by saying brains form memories from embodied events. Perception, emotion, context, action, personal history. LLMs learn statistical structure from data. And they immediately warn against calling a model component an artificial hippocampus.
Tom: Right, they say the goal is not a one-to-one comparison of anatomy, but a comparison of experimental questions.
Jane: And then they introduce Box one which defines key terms. This is really important because they're careful to say these definitions don't imply equivalent mechanisms. Lived episodic memory is biological. Parametric knowledge is LLM. Context-window information is LLM. External memory is LLM system.
Tom: So they're building a vocabulary that lets you talk about both systems without conflating them.
Jane: Exactly. And then they lay out the evidence asymmetry. Human single-neuron studies show invariant concept responses, temporal binding, rapid association formation, episode-specific coding, reactivation before recall. But these observations are rarely followed by selective manipulation of the same neurons.
Tom: Whereas rodent studies can tag, activate, silence, or reassociate engrams. But macaque and human experiments usually perturb broader circuits.
Jane: And then LLMs. They can patch activations, edit weights, steer representation directions, and rewrite external stores. All while measuring consequences throughout the system.
Tom: But here's the key caveat. The authors say this greater access doesn't mean LLMs have better memory.
Jane: Right. The edited object is usually a factual association or a transient computational state, not a lived episode embedded in perception and emotion. That distinction is explicit.
Tom: So the first page is really about setting the terms of engagement. You can compare the questions, but you can't compare the anatomy.
Jane: And they introduce Figure one which shows this asymmetry visually. Human and macaque studies give rich phenomenology but broad intervention. Rodents give selective intervention. LLMs give direct access to weights, activations, representation directions, and external stores.
Tom: So the first page is essentially the mission statement. We're going to compare experimental questions, not biological parts.
Jane: And that sets up everything else. The four shared questions, the reverse pipeline, the guardrails. It all flows from this initial framing.
Tom: I think that's why this paper is getting attention. It's not just another "AI is like a brain" paper. It's a methodological proposal.
Jane: It really is. And the authors are explicit that the productive bridge is to transfer experimental logic, not anatomical parts.
Tom: So where does that leave us? What's the big conclusion?
Jane: That's our final segment. Let's wrap it up.
Conclusion: Tom: We're wrapping up our discussion of "From Observation to Intervention: Memory in Brains and Large Language Models." Jane, what's the big takeaway for our listeners?
Jane: The big takeaway is that brains and LLMs are best compared through shared functional questions, not assumed biological equivalence. The paper argues that LLMs aren't ahead in memory itself, but in experimental access. And that access can be used to generate sharper biological hypotheses.
Tom: So instead of just watching brains and wondering, we can now intervene in models and then go test those ideas in mice, macaques, and humans.
Jane: Exactly. And the authors lay out a concrete agenda. How do you distinguish retrieval from response selection? Can you update one association without disrupting related ones? What determines whether an intervention is temporary or lasting? And which principles found in LLMs also appear in biological systems?
Tom: Those are the four questions that structure the whole paper.
Jane: And they're careful to include guardrails. An LLM result is a hypothesis generator, not evidence about the brain. Don't map brain regions to model components. Test for unintended effects. And null results are informative.
Tom: So if something works in a model but fails in a mouse, that tells you something about biological memory.
Jane: It tells you that the body, emotion, neuromodulation, and lived experience matter. And those are exactly the things LLMs don't have.
Tom: So the paper is really a call for a two-way street. Neuroscience has inspired AI for decades. Now AI can inspire neuroscience.
Jane: And the authors think this could lead to discoveries that aren't obvious from observation alone. The next frontier isn't a better metaphor between brains and models. It's a stronger experimental cycle. Identify a principle in an LLM, intervene on it, measure the result, and test whether a related process appears in biological memory.
Tom: That's a beautiful way to put it. And I think that's why this paper matters. It's not just a review. It's a proposal for how to do research differently.
Jane: And it's a proposal that respects both fields. It doesn't say LLMs are brains. It doesn't say brains are LLMs. It says we can learn from each other by asking the same questions and using the tools each system offers.
Tom: Well said, Jane. So for our listeners, if you're interested in memory, AI, or neuroscience, this paper is worth your time. "From Observation to Intervention: Memory in Brains and Large Language Models."
Jane: And we'll be back soon with another paper. Until then, keep asking good questions.
Tom: Take care, everyone.
Morteza Salehjahromi, Shayan A. Zadegan, Amgad Muneer, Jia Wu
The University of Texas MD Anderson Cancer Center · The University of Tennessee Health Science Center
q-bio.NC, cs.AI, cs.CL
Submitted: 2026-07-27
Updated: 2026-08-14
Comments: Perspective article, 11 pages, 3 figures, 1 table, and 1 key terms box. Submitted for consideration to Nature Machine Intelligence
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 66/100
Terminology
Summary
Summary
This perspective paper argues that brains and large language models (LLMs) are fundamentally different memory systems, but they can be compared through shared functional questions rather than anatomical equivalence. The authors state: Brains and LLMs are most usefully compared through shared functional questions, not assumed biological equivalence.
The core questions are: where information is represented, how partial cues retrieve broader associations, how new information is written or updated, and what changes when memory-related states are perturbed.
The paper highlights a marked asymmetry in experimental access between biological and artificial systems. Human single-neuron studies reveal invariant concept responses, temporal binding, rapid association formation, episode-specific coding, and reactivation before verbal recall,
yet these observations are rarely followed by selective manipulation of the same neurons.
Rodent engram studies can tag, activate, silence, or reassociate neuronal ensembles recruited during learning,
while macaque and human experiments usually perturb broader circuits with effects depending on target, timing, frequency, and task.
In contrast, LLM studies can patch activations, edit weights, steer representation directions, and rewrite external stores while measuring consequences throughout the system.
The authors emphasize that this greater access does not mean that LLMs possess better memory.
The edited object in LLMs is usually a factual association, transient computational state, behavioral tendency, or stored record, not a lived episode embedded in perception and emotion.
The comparison becomes useful only when that distinction is explicit.
The paper organizes the comparison around four shared questions. First, representation: biological memory is distributed across synaptic changes, neuronal ensembles, hippocampal–cortical interactions, and population-level patterns,
while LLM information may be encoded in model weights, expressed through feed-forward and attention computations, represented in residual-stream activations, maintained in the context window, or stored in external memory systems.
Both fields should study memory at several levels rather than looking for one storage location. Second, retrieval from partial cues: a partial cue can trigger the brain to recover a more complete memory through hippocampal–cortical interactions, while in LLMs a prompt shapes processing and external retrieval systems like HippoRAG, inspired by hippocampal indexing, uses an associative graph to connect and retrieve information across multiple documents.
The shared principle is pattern completion, though human recall reconstructs a previously experienced event whereas an LLM produces an answer from learned statistical associations. Third, writing and updating: human medial temporal-lobe neurons can expand selectivity after few exposures, and complementary learning systems distinguish rapid learning from slower integration; in LLMs, weights are not normally changed during use, and new information may remain in context, be added to an external store, or require training or model editing. Both fields face the stability–plasticity problem. Fourth, intervention: biological interventions range from selective engram manipulation in mice to broader stimulation in macaques and humans, while LLM interventions include ROME (persistent weight change), ActAdd (activation changes during inference), representation engineering (shifting broader directions), and RAG (changing a separable external store). The authors stress that temporary control should be distinguished from lasting change.
A key table (Table 1) summarizes intervention capabilities. Rodent studies can selectively manipulate an identified memory-related state
and make a lasting, targeted change to one association,
while human and macaque studies generally cannot. LLMs can broadly perturb systems, change valence, modify external stores, and make lasting targeted changes, but no system can delete one target memory or association without changing related knowledge,
edit one target memory or association with guaranteed zero unintended effects,
or trace the complete causal path from one representation to complex behavior or output.
The paper proposes a reverse experimental pipeline
in which LLM interventions generate hypotheses for biological testing. The authors state: "The familiar direction of influence has often been from biology to AI... The questions above motivate a complementary direction: LLM intervention → testing in the most suitable biological model → validation in a system with richer behavior → human investigation. Four questions define the research agenda:
How can memory retrieval be distinguished from later response selection, confidence, and report? Can one learned association be updated without disrupting related memories? What determines whether an intervention produces a temporary change in recall or a lasting memory update? Which principles identified in LLMs also appear in mice, macaques, and humans?"
The paper provides an example of tracing competition between old and updated associations, referencing Meng et al.'s causal tracing in GPT and Gelbard-Sagiv et al.'s human recordings. It cautions that layer order and token order represent different kinds of timing in an LLM
and should not be matched directly to brain regions. Possible future LLM findings could motivate hypotheses about whether the brain first reactivates an older memory, whether different neural populations contribute to retrieval versus response selection, and whether individual differences in updating speed predict intrusion errors or lower confidence.
The paper also proposes a shared evaluation language for memory updating,
using criteria from LLM editing research: whether an intervention changes only the intended association, generalizes across reworded prompts, preserves unrelated knowledge, lasts over time, and avoids unintended effects. These criteria could form a shared benchmark for LLM and biological memory interventions
that asks whether an intervention works, affects only the intended association, generalizes to related cues, lasts or can be reversed, causes unintended changes, and whether its effects can be traced to internal changes.
Finally, the paper lists four guardrails for the comparison: "First, an LLM result is a hypothesis generator, not evidence about the brain. Second, the comparison should focus on shared functions and experimental questions, not on matching brain regions to model components. Third, outputs alone are not enough: interventions should also be tested for unintended effects, changes to related knowledge, and how long the effect lasts. Fourth, null results are informative."
The conclusion states: "LLMs may therefore offer an experimentally tractable setting for studying problems that remain difficult in biological systems, including selective updating, persistence, reversibility, and unintended effects. This access can help generate sharper biological hypotheses. After decades in which neuroscience inspired AI, the direction can now partly reverse: discoveries in LLMs can motivate targeted experiments in mice, macaques, and humans. The next frontier is not a better metaphor between brains and models, but a stronger experimental cycle: identify a principle in an LLM, intervene on it, measure the result, and test whether a related process appears in biological memory."
Improvements for AI systems
Based on the paper, here are specific improvements I can implement in AI systems, along with what the improved system can do:
Improvement: Add an explicit, inspectable retrieval stage before response generation. The model first produces a candidate memory/association, then a separate selection stage decides whether to use it.
What it can do:
-
Distinguish between
the model retrieved the old fact
andthe model output the old fact
-
Detect when an outdated association is still active internally but suppressed at output
-
Enable intervention at the retrieval stage vs. the selection stage, with different effects
-
Provide a measurable signal for
memory competition
between old and updated associations
These improvements are directly motivated by the paper's four shared questions (representation, retrieval, updating, intervention) and its proposed reverse pipeline. They make LLM memory systems more inspectable, more controllable, and more useful for generating and testing biological hypotheses.
Sources
- Representation Engineering: A Top-Down Approach to AI Transparency
- Steering Language Models With Activation Engineering
- The AI Hippocampus: How Far are We From Human Memory?
- Memory-Augmented Transformers: A Systematic Review from Neuroscience Principles to Enhanced Model Architectures
Related papers
- BrainWave: A Brain Signal Foundation Model for Clinical Applications
- Toward Robust, Reproducible, and Widely Accessible Intracranial Speech Brain-Computer Interfaces: A Comprehensive Narrative Review of Neural Mechanisms, Hardware, Algorithms, Evaluation, Clinical Pathways and Future Directions
- CytoNet: A Foundation Model for the Human Cerebral Cortex at Cellular Resolution
- Emergence of psychopathological computations in large language models
- NeuroAI and Beyond: Bridging Between Advances in Neuroscience and Artificial Intelligence
- Attraction to hierarchical feature memory explains orientation bias