Memory by Design: Probabilistic Sequence Layers

arXiv:2605.31163 · stat.ML, cs.LG · Submitted 2026-05-29 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Memory by Design: Probabilistic Sequence Layers".

Jane: The gist The design-model framework introduces a way to derive efficient recurrent sequence maps from explicit assumptions about memory,

Tom: First, who's behind it and why it matters.

Title and authors: Tom: So we’re talking about this paper now, "Memory by Design: Probabilistic Sequence Layers." It sounds a bit dense, but really the core idea is that you can build efficient recurrent sequence maps by making explicit assumptions about how memory works.

Jane: That's right. Instead of just building a recurrence and hoping it works, they introduce this design-model framework where you separate the writing of evidence into memory from how you actually read that memory out to get an output.

Lu: It’s about this separation of the three things usually tangled up in a recurrence: the latent memory, the way you write information using Bayesian filtering, and then what produces the final activation through a readout model.

Meng: So they aren't just looking at one single way to make a layer; they’re showing how different architectures fit under this general design-model assumption.

Lalam: It sounds like they are giving us a blueprint for building these sequence models from scratch, based on probabilistic rules about memory.

The paper's summary: Tom: In terms of what they actually did, the paper introduces this design-model framework where you have a design model that writes evidence into memory using exact Bayesian filtering and then a query-dependent readout gives you a predictive distribution whose mean is the layer output.

Jane: So, to put it simply, they’ve separated the process of writing into memory from the process of reading out what’s in memory to get your final result.

Lu: They show that this separation lets them derive recurrent layers from four ingredients: an auxiliary probabilistic model for writing to internal memory, an initial belief over that memory, a readout model for querying it, and finally an output rule that converts the predictive distribution into the deterministic layer output.

Meng: That means you have more control over the dynamics because you can tweak how much uncertainty is carried forward versus how much is discarded.

Lalam: It really boils down to making sure that when you write something, it’s filtered probabilistically, and then when you read it back, that read operation is also defined separately from the writing process.

The paper's improvements: Tom: One of the big things they highlight is how this framework unifies several existing models. They show that linear attention, GLA, and Mamba-two/SSD are exact filters when you use a latent-input design model <ref:2605.31163#pg1,linear attention, GLA, and Mamba-2/SSD are exact filters>.

Jane: But they also show how other architectures like DeltaNet fit in by treating them as reductions of the Bayesian Layer's design model through a covariance reset.

Lu: They point out that Delta-rule corrections, for example, come from applying a covariance reset inside the Bayesian Layer’s design model, while additive recurrences are exact under a different auxiliary model within the same framework.

Meng: What really stands out is how restoring covariance propagation lets you get closed-form predictions for retrieval dynamics, which they then verify in controlled collision studies and other benchmarks.

Lalam: It suggests that you can get a write rule where the gain at a specific memory location depends on what memory has already resolved there, which is shown by this formula: g ss A = σ¯2A,∞ / (r squared + ¯σ2A,∞) <ref:2605.31163#pg1>.

Conclusion: Tom: So to wrap up the "Memory by Design: Probabilistic Sequence Layers" paper, it’s about deriving recurrent sequence maps from explicit assumptions about memory dynamics using a design model that separates writing and reading.

Jane: They show this separation allows for a layer that propagates both a mean state and a covariance state, meaning the gain of your write direction is shaped by the uncertainty you already have in memory.

Lu: The real contribution is having this framework derive these rules from explicit probabilistic assumptions about how information should accumulate and be retrieved, not just fitting parameters to data.

Meng: From an engineering standpoint, being able to control that uncertainty steering the writes means we can design layers that are more robust when they encounter unexpected inputs during long context retrieval.

Lalam: This paper gives us a way to build sequence models where the memory updates and the readout are truly distinct ingredients, which should make building better AI systems much more structured.

Champalimaud Research · Champalimaud Foundation

stat.ML, cs.LG

Submitted: 2026-05-29

Updated: 2026-10-07

Comments: Preprint, in submission

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 83/100

The gist: The gist The design-model framework introduces a way to derive efficient recurrent sequence maps from explicit assumptions about memory, where a design model writes evidence into memory by exact

Key concepts

Design Model
This model defines how evidence is written into an internal memory based on a fixed input prefix. It does not generate the entire input sequence likelihood but instead provides a conditional model whose Bayesian filter acts as the layer's memory update mechanism. It separates the writing process from the final output.
Bayesian Layer
This instantiation uses linear Gaussian assumptions to store both a mean state (current associations) and a covariance state (tracking how confidently each memory direction is resolved). This propagation automatically yields features like gain decay and uncertainty-steered keys, allowing the layer to adapt its memory access based on stored evidence.
Covariance Propagation
This refers to tracking the uncertainty associated with each memory direction. The framework shows that repeated writes to the same address stabilize this variance. This propagation enables an uncertainty-adaptive write rule where a key's gain depends on how much memory has already resolved that specific location.
Unifying Architectures
The framework unifies various recurrent models by showing they are either exact filters under a latent-input design model or covariance-reset reductions of the Bayesian Layer. This demonstrates a unified mathematical structure for different types of sequence modeling, including linear attention and Mamba variants.

Terminology

Summary

The gist The design-model framework introduces a way to derive efficient recurrent sequence maps from explicit assumptions about memory, where a design model writes evidence into memory by exact Bayesian filtering and a query-dependent readout produces a predictive distribution whose mean is the layer output.

The Design-Model Framework

The framework derives recurrent layers from auxiliary probabilistic design models that separate three choices usually entangled in the recurrence: the latent memory, the input-parameterized write, and the belief readout that produces the activation. A design model is not a generative model of the input sequence and defines no likelihood for x1:T. For a fixed input prefix, it defines an auxiliary conditional model whose Bayesian filter [Särkkä, 2013] becomes the layer’s memory update. This separation distinguishes recurrences that propagate uncertainty about memory from those that collapse or discard it. Formally, a recurrent sequence layer is a causal map x1:T 7→ y1:T, yt = Ft(x1:t), (1) which the framework derives from four ingredients: an auxiliary probabilistic model for writing to an internal memory, an initial belief over that memory, a readout model for querying it [Schmidhuber, 1992, Graves et al., 2014], and an output rule that converts the resulting predictive distribution into the deterministic layer output.

The Bayesian Layer and Covariance Propagation

The linearGaussian instantiation in §2.1 gives the Bayesian Layer, which propagates both a mean state, storing associations, and a covariance state, tracking how confidently each memory direction has been resolved. The layer maintains a belief over this memory where its mean stores the current associations; its covariance tracks which key directions have been resolved and therefore how future inputs should update memory. This propagation yields automatic gain decay, directional gating, and uncertainty-steered keys.

Unifying Existing Architectures

The framework unifies several sub-quadratic recurrences: linear attention, GLA, and Mamba-2/SSD are exact filters under a latent-input design model, whereas DeltaNet and related Delta-rule models are covariance-reset reductions of the Bayesian Layer’s design model. Delta-rule corrections [Schlag et al., 2021, Yang et al., 2025, Kimi Team, 2025] arise from a covariance reset within the Bayesian Layer design model, while additive recurrences [Katharopoulos et al., 2020, Yang et al., 2024a, Dao and Gu, 2024] are exact under a different auxiliary model in the same framework (Table 1).

Empirical Verification of Covariance Propagation

Restoring covariance propagation yields closed-form predictions for retrieval dynamics, which we verify empirically, and improves robustness beyond the training regime in controlled collision studies, learned associative recall under extrapolation (§4.1), and the Zoology MQAR benchmark [Arora et al., 2024a] (§4.2). The covariance geometry of §2.3 predicts that key collisions are destructive only when updates ignore directional uncertainty. Repeated writes to the same address kA drive its directional variance σ¯2t(kA) to a finite limit (42).

Performance in Long-Context Retrieval

Distilling Bayesian Layers into a pretrained 340M Gated DeltaNet improves RULER long-context retrieval over a matched-compute control, at a 2.5–2.7% held-out perplexity cost. The retrieval gains also depend on position (Figure 5a) and are strongest on multi-target retrieval.

Conclusion

The design-model framework yields a recurrence that propagates both a mean state and a covariance state, so that write direction, gain, and gating are shaped by the accumulated uncertainty over stored evidence. The propagated covariance and the design-model framework make distinct contributions: covariance gives the layer an uncertainty-adaptive write rule whose gain at a key depends on what memory has already resolved there, whereas the framework derives this rule and its fixed-gain reductions from explicit probabilistic assumptions.

Limitations

Computationally, each Bayesian head carries a grouped covariance state with G dense D × D blocks alongside its mean memory, of total recurrent size O(Dm + GD2) and per-token cost O(GD2), with quadratic working storage in D. Low-rank, sparse, or otherwise structured covariance approximations could reduce this cost and are left for future work.

Acknowledgments

M.D., H.J., and I.M.P. were supported by NIH RF1-DA056404 and by the Portuguese Recovery and Resilience Plan (PRR) through project 62 (Center for Responsible AI), and by Portuguese national funds through FCT (Fundação para a Ciência e a Tecnologia) in the context of project UIDB/04443/2020.

References

Brian D. O. Anderson and John B. Moore. Optimal Filtering. Prentice-Hall, Englewood Cliffs, N.J., 1979.

Simran Arora, Sabri Eyuboglu, Aman Timalsina, Isys Johnson, Michael Poli, James Y Zou, Atri Rudra, and Christopher Ré. Zoology: Measuring and improving recall in efficient language models. In International Conference on Learning Representations, volume 2024.

Simran Arora, Sabri Eyuboglu, Michael Zhang, Aman Timalsina, Silas Alberti, Dylan Zinsley, James Zou, Atri Rudra, and Christopher Ré. Simple linear attention language models balance the recall-throughput tradeoff. In Proceedings of the 41st International Conference on Machine Learning.

Maximilian Beck, Korbinian Pöppel, Markus Spanring, Andreas Auer, Oleksandra Prudnikova, Michael Kopp, Günter Klambauer, Johannes Brandstetter, and Sepp Hochreiter. xLSTM: Extended long short-term memory. In Advances in Neural Information Processing Systems.

Ali Behrouz, Zeman Li, Praneeth Kacham, Majid Daliri, Yuan Deng, Peilin Zhong, Meisam Razaviyayn, and Vahab Mirrokni. ATLAS: Learning to optimally memorize the context at test time.

Djohan Bonnet, Jamie Lohoff, Jan Finkbeiner, Elidona Shiqerukaj, and Emre Neftci. Learning to remember, learn, and forget in attention-based models. In International Conference on Machine Learning.

Tri Dao and Albert Gu. Transformers are SSMs: Generalized models and efficient algorithms through structured state space duality. In International Conference on Machine Learning.

Soham De, Samuel L. Smith, Anushan Fernando, Aleksandar Botev, George Cristian-Muraru, Albert Gu, Ruba Haroun, Leonard Berrada, Yutian Chen, Srivatsan Srinivasan, Guillaume Desjardins, Arnaud Doucet, David Budden,, Yee Whye Teh,, Razvan Pascanu,, Nando de Freitas,, and Caglar Gulcehre. Griffin: Mixing gated linear recurrences with local attention for efficient language models. arXiv preprint arXiv:2402.19427.

Jusen Du, Weigao Sun, Disen Lan, Jiaxi Hu, Tao Zhang, and Yu Cheng. MoM: Linear sequence modeling with mixture-of-memories. In International Conference on Learning Representations.

Alex Graves, Greg Wayne, and Ivo Danihelka. Neural Turing machines. arXiv [cs.NE], October 2014.

Improvements for AI systems

  1. We can design recurrent sequence layers that explicitly model memory dynamics via a design model, allowing for control over how evidence accumulates in memory by defining write dynamics (Bayesian filtering) separately from readout (query-dependent prediction).

  2. The Bayesian Layer propagates both a mean state and a covariance state, where the covariance tracks uncertainty over stored associations, steering writes toward uncertain directions and preserving confident memories.

  3. The framework unifies several sub-quadratic recurrences: linear attention, GLA, and Mamba-2/SSD are exact filters under a latent-input design model while DeltaNet and related Delta-rule models are covariance-reset reductions of the Bayesian Layer’s design model.

  4. We can achieve quantitative predictions for retrieval behavior by leveraging covariance propagation to derive closed-form predictions for retrieval dynamics, which we verify empirically, including predicting the steady-state write-gain decay and the cross-direction crowding rate (1−ρ2l2).

  5. The Bayesian Layer's uncertainty-aware update mechanism allows for a write rule where the gain at a key depends on what memory has already resolved there, as shown by the closed-form prediction of the scalar gain: g ss A = σ¯2A,∞ / (r squared + ¯σ2A,∞).

  6. We can improve long-context retrieval without inducing a generic language-modeling trade-off by distilling Bayesian Layers into larger pretrained models; specifically, BL gains +1.3 / +2.9 / +3.2 / +1.0 over GDN-zero at 2k / 4k / 8k / 16k in the RULER NIAH benchmark.

  7. The system can be adapted to handle complex, structured address streams by using per-column and grouped observation noise, allowing for per-column gating where each column maintains its own covariance and gain, instead of a single shared belief covariance.

Sources

Related papers