On Memory: A comparison of memory mechanisms in world models
cs.AI, cs.LG
Submitted: 2025-12-07
Updated: 2026-09-26
Comments: 10 pages, 1 figure. Published in the Transactions on Machine Learning Research
Journal ref: Transactions on Machine Learning Research (09/2026), Paper 8547
License: http://creativecommons.org/licenses/by/4.0/
The gist: World models enable agents to plan within imagined environments by predicting future states conditioned on past observations and actions.
Terminology
Abstract
World models enable agents to plan within imagined environments by predicting future states conditioned on past observations and actions. However, their ability to plan over long horizons is limited by the effective memory span of the backbone architecture. This limitation leads to perceptual drift in long rollouts, degrading the model's capacity to recall recently observed scenes. In this work, we investigate the effective memory span of transformer-based world models through an analysis of memory augmentation mechanisms. We introduce a taxonomy that distinguishes between memory encoding and memory injection mechanisms, motivating their roles in extending the world model's memory through the lens of residual stream dynamics. We evaluate twenty combinations of four encoding methods and five injection methods in the MemoryMaze environment. Using a state recall evaluation task across multiple imagination horizons, we measure the memory recall capacity of each mechanism and analyze their respective trade-offs in reconstruction quality, latent prediction error, and computational cost. We further ablate the effect of injection depth and compare the best memory-augmented vision transformer against a pure state-space model backbone. Our central finding is that the mLSTM memory encoder outperforms all alternatives in both reconstruction and latent fidelity metrics. Paired with additive injection, it exhibits the strongest recall capabilities at a moderate computational cost while matching or slightly exceeding a pure Mamba backbone. This evaluation is limited to a single environment and does not explore the effect of each mechanism on downstream task performance. We believe these research directions merit an in-depth study of their own to clearly isolate the effects, and therefore are left for future work.
Sources
- Titans: Learning to Memorize at Test Time
- ATLAS: Learning to Optimally Memorize the Context at Test Time
- Memory Layers at Scale
- A Learned Representation For Artistic Style
- World Models
- Evaluating Long-Term Memory in 3D Mazes
- Long-Context State-Space Video World Models
- Pretraining with hierarchical memories: separating long-tail and common knowledge
- StateSpaceDiffuser: Bringing Long Context to Diffusion World Models
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Highway Networks
- WorldMem: Long-term Consistent World Simulation with Memory
- DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection