Knowledge Acquisition During Pre-training? Large Language Models Learn Better With Auxiliary Views
cs.CL, cs.AI
Submitted: 2026-09-03
Updated: 2026-09-03
Comments: Accepted to Findings of EMNLP 2026
Code: https://github.com/jiosephlee/auxiliary-views-knowledge-acquisition
License: http://creativecommons.org/licenses/by/4.0/
The gist: Gaps remain in our understanding of how large language models (LLMs) acquire knowledge during pre-training.
Terminology
Abstract
Gaps remain in our understanding of how large language models (LLMs) acquire knowledge during pre-training. We posit that auxiliary views, reformulations of knowledge, are causally helpful for learning. We design controlled experiments to isolate this. First, we confirm that repetition is necessary for acquisition and clarify that paraphrasing helps only at smaller batch sizes. Second, holding the token budget fixed, allocating tokens from document repetition to auxiliary views improves learning, counterintuitively, even for factual recall. Third, the effectiveness of auxiliary views is not contingent on the strength of the teacher model that generates them. Fourth, we identify forms of knowledge, contextual and foundational, that aid learning in the presence of prior knowledge gaps. Finally, we examine how these effects manifest mechanistically via layer-wise biases and compression. Together, our findings suggest that auxiliary representations of knowledge, which arise naturally in large pre-training corpora, are a key factor in the success of pre-training and offer a plausible explanation for why data diversity matters.
Sources
- Intern-S1: A Scientific Multimodal Foundation Model
- SaulLM-7B: A pioneering Large Language Model for Law
- Textbooks Are All You Need
- Scaling Laws and Interpretability of Learning from Repeated Data
- Meta-Learning Online Adaptation of Language Models
- Towards Continual Knowledge Learning of Language Models
- Scaling Laws for Neural Language Models
- Deduplicating Training Data Makes Language Models Better
- 2 OLMo 2 Furious
- Reuse, Don't Retrain: A Recipe for Continued Pretraining of Language Models
- Infini-gram: Scaling Unbounded n-gram Language Models to a Trillion Tokens
- MedGemma Technical Report
- Galactica: A Large Language Model for Science
- TxGemma: Efficient and Agentic LLMs for Therapeutics
- Towards Objective Fine-tuning: How LLMs' Prior Knowledge Causes Potential Poor Calibration?
- OPT: Open Pre-trained Transformer Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering