LMEnt: A Suite for Analyzing Knowledge in Language Models from Pretraining Data to Representations
cs.CL
Submitted: 2025-09-03
Updated: 2026-09-17
Comments: Accepted to Transactions of the Association for Computational Linguistics (TACL) 2026
License: http://creativecommons.org/licenses/by/4.0/
The gist: Language models (LMs) increasingly drive real-world applications that require world knowledge.
Terminology
Abstract
Language models (LMs) increasingly drive real-world applications that require world knowledge. However, the internal processes through which models turn data into representations of knowledge and beliefs about the world are poorly understood. To facilitate such studies, we present LMEnt, a suite including (1) a knowledge-rich pretraining corpus, fully annotated with entity mentions based on Wikipedia, (2) an entity-based retrieval method over pretraining data that outperforms existing tools by as much as 80.4%, and (3) 12 pretrained LMs with up to 1B parameters and 4K intermediate checkpoints, with comparable performance to popular open-source models on knowledge tasks. Together, these resources provide a controlled environment for analyzing connections between entity mentions in pretraining data and downstream performance. We show the utility of LMEnt by studying knowledge acquisition over training, finding that entity co-occurrence and mention forms-which are difficult to study with existing tools-affect learning trends. Moreover, as LMs form stronger associations between entities, their facts are harder to edit in-context, whereas inconsistencies in model predictions over training are indicative of editing success. We release LMEnt to support studies of knowledge in LMs, including knowledge representations, plasticity, editing, attribution, hallucinations, and learning dynamics.
Sources
- TinyStories: How Small Can Language Models Be and Still Speak Coherent English?
- A Survey on LLM-as-a-Judge
- NeMo: a toolkit for building AI applications using Neural Modules
- RoBERTa: A Robustly Optimized BERT Pretraining Approach
- 2 OLMo 2 Furious
- LLaMA: Open and Efficient Foundation Language Models
- Llama 2: Open Foundation and Fine-Tuned Chat Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering