Can a Language Model Learn Facts Continually in Its Weights?
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Can a Language Model Learn Facts Continually in Its Weights?".
Jane: Continual learning promises a language model that keeps acquiring knowledge after training, with each new fact written into its weights.
Tom: First, who's behind it and why it matters.
Paper summary: Tom: So we're talking about this paper here titled "Can a Language Model Learn Facts Continually in Its Weights?" and it’s Charles O’Neill, and it looks like they are looking at whether an AI can keep learning new facts after the initial training is done by actually writing those facts into its weights.
Jane: Right, so the main idea here is that continual learning promises a language model that keeps acquiring knowledge after training, with each new fact written into its weights. But the paper immediately raises this big question about whether those weight writes can actually support the accumulation of knowledge over time.
Lu: It's really interesting because they follow invented facts written into Qwen3 models from creation through sequences of twenty to one hundred later writes, using held-out questions of five types, with the original model given the fact in its prompt as the reference. They’re testing exactly what happens when you try to store that new information in two different ways: either written into weights or just placed in context.
Meng: So they’re looking at how these later writes preserve knowledge, and what remains after questions fail, which is important for figuring out if this learning actually sticks or if it just gets lost.
Lalam: I see the core experiment involves tracking invented facts through sequences of twenty to one hundred later writes using held-out questions of five types. That’s a pretty structured way to test persistence.
Tom: Exactly, and what they find is that the breadth of the training data determines the kind of knowledge created, because bare-statement training produces recitation, while diverse restatements reduce the recitation-to-use gap from twenty-seven point four to five point four points without showing the model a conclusion <ref:2607.11020#pg0,bare-statement training produces recitation, while diverse restatements reduce the recitation-to>.
Jane: That difference carries into later writes because after twenty sequential writes, bare-statement facts retain only one percent accuracy, but facts written from broad study data retain forty-six percent <ref:2607.11020#pg0,after twenty sequential writes, bare-statement facts retain>. That’s a huge gap in how well those models hold onto the new stuff.
Lu: They also find that facts can be behaviourally forgotten without being erased; forgotten facts keep most of the log-probability added by their write, and under bare-statement training, they have about seventy percent wrong answers about them <ref:2607.11020#pg0,also find that facts can be behaviourally forgotten without being erased; forgotten>.
Meng: That suggests that even if the fact isn't perfectly recalled, it still has some residual influence on how the model behaves. So what does this mean practically for someone who just uses an AI every day?
Paper summary: Lalam: It means we need to figure out how to make sure that whatever new information we feed it doesn't just get lost in the noise of later updates.
Tom: The researchers then separate three requirements for continual writing, which are each pretty crucial: first, each write must create usable knowledge; second, general abilities must survive accumulation; and third, earlier facts must remain reachable.
Jane: And for general abilities to survive accumulation, they suggest using broad data creates usable knowledge and that a frozen original model or a penalty on local drift can protect those general abilities.
Lu: Reachability remains unsolved across supervised fine-tuning and both offline and online distillation; in-context use erodes no faster than general ability.
Meng: That’s the part where I wonder about practical impact, because reachability still isn't fully solved across different methods of updating the model.
Lalam: And when you think about holding a new fact, the paper points out that context only lasts for the life of the prompt, so context is not a reliable channel when facts must be composed or outlive later writes.
Tom: The text says that the reliable channel when facts must be composed or outlive later writes is context rather than the weights. It seems like we have to keep thinking about how we store that information if we want it to last longer than just a single prompt.
Jane: So, how do you bridge this gap between storing things in weights versus using context for newer facts? That’s what they are trying to figure out.
Lu: The paper suggests that the write stores content, while the context supplies its address at use time, which provides a general-purpose handle for surviving further writing and even working for forgotten facts as well as fresh ones.
Meng: So if we use context to store the address, does that mean we can actually pull up old facts later without retraining or constantly feeding the whole history back in?
Lalam: The paper suggests that supplying the same content in context provides a general-purpose handle for joint use, it survives further writing, and it works for forgotten facts as well as fresh ones. It’s about using context to supply an address at use time.
Tom: That brings us to the practical guidance they give: you need broad data for each write, you need to avoid bare-statement training when later updates must preserve existing knowledge, and you should distill against a frozen copy of the original model.
Jane: So what’s their final thought on where continual learning belongs in terms of system design? Do they have a specific channel in mind that handles both storing content and creating an address?
Paper summary: Lu: They conclude that continual learning belongs in a channel that creates an address as well as storing content, which is the key idea tying everything together.
Meng: That sounds like we’re looking for some kind of architectural layer beyond just updating the weights or just using context windows to solve this problem.
Lalam: It really points toward needing a system that handles both writing and addressing simultaneously if we want to achieve what they are aiming for in continual learning.
Tom: So, moving on from the technical details, let’s look at the conclusion of this paper, "Can a Language Model Learn Facts Continually in Its Weights?" and what it all means for us.
Jane: The authors are Charles O’Neill and they are asking this fundamental question about how language models handle new information over time when learning is happening within their weights.
Lu: The implication is that whether weight writes can support accumulation remains undecided, but the research shows that context offers a reliable channel for facts if you need them to outlive later writes.
Meng: So for the people listening just tuning in, it means we have to be really careful about what kind of data we feed these models during updates if we want to keep useful knowledge instead of just adding noise.
Lalam: The paper shows that the way you structure your learning—whether you’re using bare statements or diverse restatements—has a huge effect on how well those facts stick and are used later.
Tom: Exactly, it’s not just about adding more data; it’s about how that data is presented to the model so it can actually use what it learned before.
Jane: And the paper highlights that we need to find a way to create an address for facts beyond just the specific questions used during training, which is a big missing piece in how models handle recall.
Lu: It suggests that building this mechanism for addressing content, alongside storing it, is the direction forward for making these systems truly capable of continual learning.
Meng: From an engineering standpoint, it tells us that if we want to build systems that learn persistently, we need to design a channel specifically meant to manage both the content itself and a way to look up that content later.
Lalam: It’s about creating a channel where the write stores the content, but then something else supplies its address at use time so it can work for forgotten facts as well as fresh ones.
Tom: That’s our wrap-up on this paper, "Can a Language Model Learn Facts Continually in Its Weights?" and it points toward needing a better way to manage memory in these models.
Conclusion: Tom: So we're wrapping up this look at "Can a Language Model Learn Facts Continually in Its Weights?"
Jane: It really boils down to whether you can keep adding new facts to an AI without it just forgetting what it already knows.
Lu: The authors are Charles O’Neill and Yuanzhi Li. They're asking if we can actually get the AI to learn things while updating its core weights after the initial training is done.
Meng: So, essentially, they're testing if those weight updates can support long-term knowledge accumulation in a language model.
Tom: Right, and what they find is that it’s not a simple yes or no answer on whether this works smoothly with weights.
Jane: It seems the main takeaway is that the way you structure the learning—the kind of facts you feed it—really matters for how well those new bits stick around later.
Lu: They found that using broad, diverse data for each update makes a huge difference in retention, moving accuracy from one percent to forty-six percent after twenty writes.
Meng: That's a big jump; so the quality of the input data is super important for keeping those learned facts alive.
Tom: It also points out that context might be a better way to hold onto new facts if you need them to last longer than just the current prompt window.
Jane: So, it seems we have two main ways: trying to embed the knowledge in the weights versus using context as a separate address system for newer things.
Lu: They are looking for that specific mechanism—a way to store content and also provide an address so you can find old facts later without retraining everything.
Meng: That’s the practical hurdle; how do we build that kind of memory layer in the AI architecture?
Tom: That’s exactly where we need to look next, because figuring out this addressing system is key to making true continual learning happen.
Charles O’Neill
cs.CL, cs.LG
Submitted: 2026-07-13
Updated: 2026-10-05
Code: https://github.com/basetenlabs/cortex
Importance score: 79/100
The gist: Continual learning promises a language model that keeps acquiring knowledge after training, with each new fact written into its weights.
Key concepts
- Continual Learning in Weights
- This refers to the idea of updating a language model's core knowledge (weights) with new information after its initial training is complete. The study tests whether this process can successfully accumulate facts without losing old ones or causing performance degradation.
- Bare-Statement Training vs. Diverse Restatements
- The quality of the initial data matters significantly. Training on simple, bare statements leads to poor knowledge retention over time. Conversely, using diverse restatements of facts helps reduce the gap between recitation and actual use, leading to much better preservation of learned knowledge in later updates.
- Context vs. Weights for Storage
- The paper distinguishes between storing new information in the model's weights versus keeping it in the prompt context. While weights are permanent storage, context is limited to the current prompt's life. The study suggests that context is a more reliable channel for holding facts that must survive later writes.
- Address for Facts
- This concept refers to a general method or handle that allows a model to retrieve any stored fact, regardless of which specific questions initially led it there. Providing this 'address' through context can support joint use and help preserve facts against forgetting.
Terminology
Summary
Continual learning promises a language model that keeps acquiring knowledge after training, with each new fact written into its weights. Whether weight writes can support accumulation remains undecided.<ref:2607.11020#pg2> The gist: later writes redirect the questions that reached it, and context remains usable through the same writes.<ref:2607.11020#pg6>
How it works
The research investigates whether language models can learn facts continually by writing them into their weights after initial training, examining what kind of knowledge is created, how later writes preserve it, and what remains after questions fail.<ref:2607.11020#pg2> The core experiment involves tracking invented facts through sequences of twenty to one hundred later writes using held-out questions of five types.<ref:2607.11020#pg5>
The breadth of the training data determines the kind of knowledge created; bare-statement training produces recitation, while diverse restatements reduce the recitation-to-use gap from 27.4 to 5.4 points without showing the model a conclusion.<ref:2607.11020#pg4> This difference carries into later writes: after twenty sequential writes, bare-statement facts retain 1% accuracy while facts written from broad study data retain 46%.<ref:2607.11020#pg4>
How it works
The results separate three requirements for continual writing: each write must create usable knowledge, general abilities must survive accumulation, and earlier facts must remain reachable.<ref:2607.11020#pg6> Broad data creates usable knowledge, and a frozen original model or a penalty on local drift can protect general abilities.<ref:2607.11020#pg6> Reachability remains unsolved across supervised fine-tuning and both offline and online distillation; in-context use erodes no faster than general ability.<ref:2607.11020#pg6>
How it works
The model can hold a new fact in its context or its weights, but context only lasts for the life of the prompt.<ref:2607.11020#pg4> The reliable channel when facts must be composed or outlive later writes is context rather than the weights.<ref:2607.11020#pg6>
How it works
The incoming write causes interference, which resists local control.<ref:2607.11020#pg6> The reliable channel is context rather than the weights when facts must be composed or survive later writes.<ref:2607.11020#pg6>
How it works
The incoming write effect is clear, while the stored-write effect is not.<ref:2607.11020#pg10> The change in retention after fifteen later writes, by later-write method, with fact-clustered 95% intervals shows that the stored-write main effect changes under reasonable definitions of successful writing, while the later-write effect does not.<ref:2607.11020#pg18>
How it works
The model can answer the single-fact questions used during training, yet cannot reliably state all of its written facts on demand or bring two of them into a joint computation.<ref:2607.11020#pg10> The missing capability is what we mean by an address: a general handle for reaching a fact beyond the questions that happen to route to it.<ref:2607.11020#pg10>
How it works
Supplying the same content in context provides a general-purpose handle: it supports joint use, survives further writing, and works for forgotten facts as well as fresh ones.<ref:2607.11020#pg10> The write stores content, while the context supplies its address at use time.<ref:2607.11020#pg10>
How it works
The practical guidance is correspondingly narrow: Use broad data for each write, avoid bare-statement training when later updates must preserve existing knowledge, and distill against a frozen copy of the original model.<ref:2607.11020#pg22> Continual learning belongs in a channel that creates an address as well as storing content.<ref:2607.11020#pg22>
How it works
The experiments reject the early-predictor, transformation-necessity, and gradient-projection hypotheses under their protocols. The boundary is accumulation, not all writing.
REFERENCES
Zeyuan Allen-Zhu and Yuanzhi Li. Physics of language models: Part 3.1, knowledge storage and extraction. arXiv preprint arXiv:2309.14316, 2024.<ref:2607.11020#pg9>
Michael C. Anderson, Robert A. Bjork, and Elizabeth L. Bjork. Remembering can cause forgetting: Retrieval dynamics in long-term memory. Journal of Experimental Psychology: Learning, Memory, and Cognition, 20(5):1063–1087, 1994.<ref:2607.11020#pg9>
Mikita Balesni, Tomasz Korbak, and Owain Evans. Lessons from studying two-hop latent reasoning. arXiv preprint arXiv:2411.16353, 2024.<ref:2607.11020#pg9>
Vincent-Pierre Berges, Barlas Oguz, Daniel Haziza, Wen-tau Yih, Luke Zettlemoyer, and Gargi Ghosh. Memory layers at scale. arXiv preprint arXiv:2412.09764, 2024.<ref:2607.11020#pg9>
Lukas Berglund, Meg Tong, Max Kaufmann, Mikita Balesni, Asa Cooper Stickland, Tomasz Korbak, and Owain Evans. The reversal curse: LLMs trained on “A is B” fail to learn “B is A”. arXiv preprint arXiv:2309.12288, 2024.<ref:2607.11020#pg9>
Dan Biderman, Jacob Portes, et al. LoRA learns less and forgets less. Transactions on Machine Learning Research, 2024.<ref:2607.11020#pg9>
Nitay Calderon, Eyal Ben-David, Zorik Gekhman, Eran Ofek, and Gal Yona. Empty shelves or lost keys? recall is the bottleneck for parametric factuality. arXiv preprint arXiv:2602.14080, 2026.<ref:2607.11020#pg9>
Stephanie C. Y. Chan, Ishita Dasgupta, Junkyung Kim, Dharshan Kumaran, Andrew K. Lampinen, and Felix Hill. Transformers generalize differently from information stored in context vs in weights. arXiv preprint arXiv:2210.05675, 2022.<ref:2607.11020#pg9>
Roi Cohen, Eden Biran, Ori Yoran, Amir Globerson, and Mor Geva. Evaluating the ripple effects of knowledge editing in language models. Transactions of the Association for Computational Linguistics, 12, 2024.<ref:2607.11020#pg9>
Kevin Lu and Thinking Machines Lab. On-policy distillation, 2025.<ref:2607.11020#pg9>
John A. McGeoch. Forgetting and the law of disuse. Psychological Review, 39(4):352–370, 1932.<ref:2607.11020#pg9>
Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov. Locating and editing factual associations in GPT. In NeurIPS, 2022.<ref:2607.11020#pg9>
Kevin Meng, Arnab Sen Sharma, Alex Andonian, Yonatan Belinkov, and David Bau. Mass-editing memory in a transformer. In ICLR, 2023.<ref:2607.11020#pg9>
Oded Ovadia, Menachem Brief, Moshik Mishaeli, and Oren Elisha. Fine-tuning or retrieval? comparing knowledge injection in LLMs. In EMNLP, 2024.<ref:2607.11020#pg9>
Shankar Padmanabhan, Yasumasa Onoe, Michael J. Q. Zhang, Greg Durrett, and Eunsol Choi. Propagating knowledge updates to lms through distillation. In NeurIPS, 2023.<ref:2607.
Improvements for AI systems
-
Bold knowledge creation through broad data training: Implement
study training
by using24 generated paraphrases, question–answer pairs, worked implications, and contrasts with the default,
which is shown to reduce the recitation-to-use gap from 27.4 to 5.4 points without showing the model a conclusion. -
Preserve knowledge retention through data diversity: Utilize
study training
overbare-statement training
when later writes must preserve existing knowledge, as facts written from broad study data retain46%
accuracy after twenty sequential writes compared to only 1% for bare-statement facts. -
Ensure fact reachability through context: For facts that must be composed or survive later writes, rely on context rather than weights; this allows the model to
recover to 77–80% on its questions
when a forgotten study fact is supplied in the prompt, whereas a forgotten study fact in weights fails entirely. -
Prevent catastrophic forgetting via teacher stability: When accumulating knowledge through distillation, use a frozen teacher instead of sequential distillation against accumulated writes; this approach
preserves general abilities
and avoids thedrifting teacher
that causes capability loss and retention collapse. -
Enhance factual retrieval with context handling: Design systems to use context as a
general-purpose handle,
as it allows forjoint use, survives further writing, and works for forgotten facts as well as fresh ones.
Sources
- Physics of Language Models: Part 3.1, Knowledge Storage and Extraction
- Lessons from Studying Two-Hop Latent Reasoning
- Memory Layers at Scale
- The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"
- Empty Shelves or Lost Keys? Recall Is the Bottleneck for Parametric Factuality
- Transformers generalize differently from information stored in context vs in weights
- Towards Mechanistically Understanding Why Memorized Knowledge Fails to Generalize in Large Language Model Finetuning
- Cartridges: Lightweight and general-purpose long context representations via self-study
- One Mask to Rule Them All: On Hidden Facts after Editing and How to Find Them
- FictionalQA: A Dataset for Studying Memorization and Knowledge Acquisition
- On the generalization of language models from in-context learning and finetuning: a controlled study
- Memorization vs. Reasoning: Updating LLMs with New Knowledge
- Narrow Finetuning Leaves Clearly Readable Traces in Activation Differences
- How much do language models memorize?
- $\textit{New News}$: System-2 Fine-tuning for Robust Integration of New Knowledge
- RL's Razor: Why Online Reinforcement Learning Forgets Less
- Self-Distillation Enables Continual Learning
- Believe It or Not: How Deeply do LLMs Believe Implanted Facts?
- Learning by Distilling Context
- How new data permeates LLM knowledge and how to dilute it
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering