Characterizing Narrative Content in Web-scale LLM Pretraining Data
cs.CL
Submitted: 2026-06-17
Updated: 2026-09-09
Comments: 8.5 pages of main content, 41 total pages. 30 figures
Code: https://github.com/johnsont4/narratives_in_
License: http://creativecommons.org/licenses/by/4.0/
The gist: The narrative composition of web-scale LLM pretraining corpora remains largely unexplored, even though narrative is a fundamental mode of human communication.
Terminology
Abstract
The narrative composition of web-scale LLM pretraining corpora remains largely unexplored, even though narrative is a fundamental mode of human communication. We present the first fine-grained study of narrative features in Dolma, a 3-trillion-token open pretraining corpus. Drawing on narrative theory, we design a framework spanning three core narrative elements (agency, setting, and events) operationalized as 11 interpretable dimensions. After curating and hand-annotating a diverse set of 400 passages, we create an LLM-labeled dataset of 25K passages, and finally, we finetune and validate NarraBERT, two RoBERTa-based models for fine-grained narrative prediction. We apply NarraBERT to 13M passages, resulting in a new dataset, NarraDolma. We find that narrative structure is measurable at scale across extremely heterogeneous data and narrative qualities are unequally distributed across pretraining sources, topics, and formats in ways that current data curation practices neither measure nor account for. Our framework, dataset, and analyses provide a foundation for understanding how narrative qualities are distributed in LLM pretraining data.
Sources
- The Llama 3 Herd of Models
- Scaling Language Models: Methods, Analysis & Insights from Training Gopher
- Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research
- DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering