Entity tracking emerges in sub-billion parameter language models and exceeds human performance in naturalistic narratives
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "Entity tracking emerges in sub-billion parameter language models and exceeds human performance in naturalistic narratives".
Tom: This research investigates whether language models (LMs) possess genuine language understanding by evaluating their capacity for entity tracking across naturalistic narratives, comparing their performance against human comprehension.
Jane: First, who's behind it and why it matters.
Title and authors: Tom: We're starting with the title of this paper, "Entity tracking emerges in sub-billion parameter language models and exceeds human performance in naturalistic narratives." It really highlights how much this research is focusing on, which is entity tracking across natural stories.
Jane: And the authors are Karolina Drozd˙ z˙ one and Micha Heilbron2,3 from the IDEAS Research Institute at Warsaw, Max Planck Institute for Psycholinguistics in Nijmegen, Netherlands. Their background suggests a really strong combination of cognitive science and advanced language modeling.
Lu: Their work seems to be bridging that gap between what we see in massive models and what humans actually do when they read a story; it’s a fascinating intersection.
Meng: I wonder how their specific evaluation methods compared these models against human performance in such complex, naturalistic settings.
Lalam: It's interesting because it challenges the idea that tracking requires these huge models to even exist at all; if smaller ones can do this, the architecture itself might be more important than just raw size.
The paper's summary: Tom: Okay, so what's the core finding? In "Entity tracking emerges in sub-billion parameter language models and exceeds human performance in naturalistic narratives," the paper summarizes that entity tracking is a core part of understanding language, meaning knowing where things are and how they move.
Jane: The key summary point is that they tested this across five levels of situational complexity for humans, finding that tracking accuracy drops when the situation gets more complex, not just because the story gets longer.
Lu: And then in language models, they found something surprising: human-level entity tracking is already present at four hundred ten million parameters. This is well below the multi-billion parameter models that were previously thought to be necessary for this task.
Meng: That number—four hundred ten million parameters—is what really catches my eye; it suggests a much lower threshold for this specific type of capability than we've been led to believe.
Lalam: This summary really points toward a more scalable approach where the ability to track things isn't entirely dependent on just having a lot of raw parameter count.
The paper's improvements: Tom: The paper also discusses how they approached the problem, suggesting that they looked at prior work which showed tracking mechanisms exist in Transformers, but not necessarily in natural language pretraining. They even identified circuits implementing entity tracking in base language models through some fine-tuning.
Jane: And one of their key methodological improvements was testing across three different types of narratives: standard objects, pseudowords, and semantically improbable objects to make sure the model wasn't just memorizing word sequences.
Lu: That part about robustness across semantic content is crucial because it shows that the tracking operates over the discourse structure itself rather than just relying on familiar words or lexical patterns.
Meng: From an engineering view, proving robustness against pseudowords means the system isn't brittle; it’s actually building a model of relationships instead of just pattern matching specific vocabulary items.
Lalam: That robustness is what makes this capability valuable in real-world applications; if the AI can track "rare kidney" the same way it tracks "car," that means its understanding is much more flexible.
Conclusion: Tom: So to wrap things up on this paper, we have to say that entity tracking is a fundamental component of comprehension, and this study shows it emerges at scales far smaller than we expected. It really suggests that the prior idea that massive parameter counts are the only path to this skill might be misleading.
Jane: Exactly; they concluded that at sufficient scale language models can actually perform better than humans in this domain because they're building those underlying situation models more effectively.
Lu: I think the implication is that we should start looking at smaller, open-weight base models with a new lens for assessing their capacity for these deep comprehension skills.
Meng: Practically speaking, this means we can potentially develop much lighter and faster AI systems that still maintain high-fidelity narrative understanding without needing to rely on huge, specialized pretraining datasets.
Lalam: I think the biggest implication is that we can start designing AI agents whose core competency isn't just answering questions but maintaining a coherent world model during long, unfolding interactions.
Tom: That’s a lot to digest about the "Entity tracking emerges in sub-billion parameter language models and exceeds human performance in naturalistic narratives" paper. We'll be right back after the break with more on these exciting developments.
Karolina Drozḋ ż, Micha Heilbron
IDEAS Research Institute · Max Planck Institute for Psycholinguistics · University of Amsterdam Brain and Cognition
cs.CL
Submitted: 2026-06-04
Updated: 2026-09-29
Importance score: 83/100
The gist: This research investigates whether language models (LMs) possess genuine language understanding by evaluating their capacity for entity tracking across naturalistic narratives, comparing their
Key concepts
- Entity Tracking
- This is a core part of understanding language, meaning knowing where things are and how they move within a story. The paper tests this ability across different levels of situational complexity to evaluate how well models understand the relationships between entities in natural narratives.
- Sub-billion Parameter Language Models
- These are language models with parameter counts significantly smaller than the massive multi-billion parameter models previously thought required for advanced understanding. The research found that entity tracking emerges at scales well below these larger models, suggesting a lower threshold for this capability.
- Robustness Across Semantic Content
- This refers to testing the model's ability to track entities when presented with different types of objects, such as standard objects, pseudowords, or semantically improbable ones. This testing proves the tracking operates over the story's discourse structure rather than just memorizing specific words.
Terminology
Summary
This research investigates whether language models (LMs) possess genuine language understanding by evaluating their capacity for entity tracking across naturalistic narratives, comparing their performance against human comprehension. The study challenges prior assumptions about the necessary scale and data requirements for such abilities, demonstrating that robust entity tracking emerges at surprisingly small model scales and that contemporary LMs can exceed human performance in this domain.
Human Performance and Cognitive Complexity
The study first established the baseline for human entity tracking by testing participants on short narratives across five levels of situational complexity (C1 to C5). The findings revealed a specific cognitive cost associated with increasing narrative difficulty: entity tracking degrades specifically with narrative complexity, not narrative length.
Specifically, in the explicit task, accuracy decreased from 87.2% at C1 to 65.0% at C4.
Crucially, controlling for recency effects by varying the position of the target object within narratives showed that "the complexity effect remained significant (b = −0.53, SE = 0.10, z = −5.45, p <.001), indicating that performance is limited by
the complexity of the situation model itself — the number of entities and relations being tracked — not by surface-level recency."
Model Scaling and Emergence of Tracking
The paper evaluates two fully open model families, Pythia (70M–12B) and OLMo 2 (1B–32B), alongside larger contemporary models like Llama 3.3 and Qwen 2.5, against human performance. The key discovery is that human-level entity tracking is already present at 410 million parameters already,
which is far smaller than the multi-billion parameter, code-specialised models identified by prior work.
Furthermore, robust, human-level tracking is present at 410 million parameters already
and improves predictably with scale. The effect of complexity diminishes as models grow, and tracking completely disappearing for contemporary models at 70B scale.
Instruction Tuning Effects
The study examined the impact of instruction tuning on entity tracking capabilities. The results showed a clear divergence between explicit and implicit performance: instruction tuning selectively improves explicit but not implicit entity tracking.
Specifically, instruction tuning substantially improved explicit task performance – with gains up to +52.4 percentage points – consistent with the task’s demand,
whereas it did not improve implicit performance, and in some cases slightly reduced it (e.g., OLMo 2 1B: 88.5% to 78.0%).
This suggests that instruction tuning enhances the ability to express tracking in a question-answering format without deepening the underlying capability.
Robustness Across Semantic Content
To ensure that model performance reflects genuine tracking rather than reliance on lexical association, models were tested on narratives containing three template conditions: standard objects, pseudowords, and semantically improbable objects. The results demonstrated that Performance was consistent across conditions,
with accuracy remaining within 3 percentage points of the standard template across all model families and evaluation formats.
This indicates that entity tracking operates over discourse structure rather than lexical familiarity,
as models registered the increased surprisal from pseudoword or improbable object narratives.
Conclusion on Model Requirements
Collectively, the results demonstrate that entity tracking – a core component of genuine language comprehension – emerges with model scale, at parameter counts far below those previously associated with this ability.
The paper concludes that prior claims suggesting high parameter counts or massive code pretraining are driven by the procedural demands of their task rather than fundamental limits on entity tracking,
and that at sufficient scale language models can far exceed human performance.
The study also provides a new evaluation paradigm for studying naturalistic entity tracking in small, fully open base models.
Improvements for AI systems
Here are the specific improvements that can be made to AI systems based on this research, categorized by capability:
)Entity Tracking in Naturalistic Narratives (Core Improvement)
The primary improvement is enabling language models (LMs) to perform genuine, human-like entity tracking across complex discourse structures in natural language, moving beyond superficial pattern matching. This capability allows the AI to maintain and update a coherent internal state model of the narrative world as it unfolds.
(Specific System Capabilities)
-
[State Maintenance and Inference]: The improved system can accurately infer unstated relationships, locations, and states of entities based on described events (e.g., inferring that a box is empty after an item is removed from it), demonstrating a genuine understanding of discourse context rather than just memorized sequences.
-
[Complexity-Dependent Reasoning]: The AI's performance in tracking entities will degrade predictably with increased situational complexity (number of objects, locations, and movements). This allows for the design of systems that can dynamically adjust their internal modeling resources based on the narrative density they are processing.
-
[Scale-Dependent Emergence]: The system's capacity for robust entity tracking emerges at significantly smaller parameter scales (around 410 million parameters) than previously reported, rather than requiring massive, code-specialized models (e.g., 13B+ parameters with specialized tokens). This suggests that fundamental representational capacity for discourse tracking is present in smaller, more general-purpose language models.
-
[Robustness to Lexical Manipulation]: The system will not rely on superficial lexical associations or distributional priors (like co-occurrence of words). It will maintain accurate tracking even when presented with semantically improbable objects (e.g.,
rare kidney
) or pseudowords, proving that it tracks the underlying discourse structure rather than surface-level vocabulary. -
[Instruction Tuning for Explicit Output]: When instruction-tuned, the system can reliably perform explicit entity tracking tasks (like answering specific location questions) with high accuracy, suggesting that alignment mechanisms effectively map internal tracking states to desired output formats (e.g., single-word location responses).
(Specific Application Areas)
-
[Advanced Narrative Comprehension]: Creating AI agents capable of maintaining complex
world models
during long, unstructured storytelling or dialogue, allowing them to answer deep questions about the narrative's state without needing explicit grounding in a database. -
[Reduced Data/Parameter Requirements for State Tracking]: Developing more efficient and smaller language models that possess human-level entity tracking capabilities, potentially leading to faster inference and reduced computational costs compared to current state-of-the-art approaches reliant on massive code pretraining.
-
[Generalization Across Domains]: Since the tracking mechanism appears robust across different object types (standard, pseudoword, improbable), these systems can be deployed in diverse, novel textual domains without needing task-specific fine-tuning for every new vocabulary set.
-
[Improved Alignment and Fine-Tuning Strategies]: Understanding that instruction tuning improves explicit output but not implicit tracking reveals a precise strategy: fine-tuning should focus on enhancing the model's ability to articulate its internal tracking state in a structured, answerable format, rather than trying to fundamentally deepen the underlying cognitive mechanism.
Sources
- To Code, or Not To Code? Exploring Impact of Code in Pre-training
- Code Pretraining Improves Entity Tracking Abilities of Language Models
- Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic Task
- The Llama 3 Herd of Models
- Fine-Tuning Enhances Existing Mechanisms: A Case Study on Entity Tracking
- Qwen2.5 Technical Report
- 2 OLMo 2 Furious
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering