Time-Incremental Continued Pretraining of LLMs: Knowledge Updates Without Catastrophic Forgetting
cs.CL
Submitted: 2026-09-20
Updated: 2026-09-20
Comments: Preprint
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large language models (LLMs) drift out of date the moment their pretraining ends, yet retraining from scratch is prohibitively expensive.
Terminology
Abstract
Large language models (LLMs) drift out of date the moment their pretraining ends, yet retraining from scratch is prohibitively expensive. Continued pretraining (CPT) is the natural remedy, but it is typically evaluated through a continual learning lens that assumes disjoint data streams. This is a poor fit for time-incremental updates on web-scale crawls, where successive snapshots share substantial URL overlap by design. We study time-incremental CPT in this realistic regime: continued pretraining on FineWeb-Edu dumps drawn strictly from after each model's knowledge cutoff, evaluated across six open-weight models spanning three families (OLMo2, Llama-3.1/3.2, Gemma-3-1B) and four parameter scales (1B-3B-7B-8B). We organize our findings around four practical questions. (i) Is knowledge acquired? Yes, but heterogeneously, and without catastrophic forgetting: five of six models also improve on pre-cutoff factual recall, and the gains track pretraining saturation (driven primarily by token budget per parameter). (ii) What does it cost? Almost nothing: the macro-average across a thirteen-task suite stays within 0.01 of the base for every model. (iii) What is the recipe? Data quality dominates quantity (a curated 6B-token slice matches a broader 40B one); the optima for knowledge acquisition and general capability are separated by roughly an order of magnitude in learning rate; and LoRA at sufficient rank matches full CPT. (iv) Does it survive deployment? CPT gains transfer through SFT, while DPO's effect is family-dependent. Together, these results paint a more optimistic picture of time-incremental CPT than the prior continual learning literature suggests.
Sources
- MEDITRON-70B: Scaling Medical Pretraining for Large Language Models
- Training Verifiers to Solve Math Word Problems
- Gemma 3 Technical Report
- The Llama 3 Herd of Models
- Continual Pre-Training of Large Language Models: How to (re)warm your model?
- Don't Stop Pretraining: Adapt Language Models to Domains and Tasks
- Code Llama: Open Foundation Models for Code
- Progressive Neural Networks
- MIRAI: Evaluating LLM Agents for Event Forecasting
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering