On the Importance of Gating: Memorization vs. In-Context Learning in State Space Models
cs.LG, cs.AI
Submitted: 2026-09-15
Updated: 2026-09-15
Comments: 25 pages, 5 figures
License: http://creativecommons.org/licenses/by/4.0/
The gist: State Space Models (SSMs) have emerged as a compelling alternative to Transformers, enabling sequence modeling with constant memory and linear compute.
Terminology
Abstract
State Space Models (SSMs) have emerged as a compelling alternative to Transformers, enabling sequence modeling with constant memory and linear compute. Although SSMs exhibit reasonable performance and favorable computational characteristics, they continue to lag behind Transformers on tasks that require in-context learning and precise retrieval, slowing their adoption for large-scale language modeling. In this work, we demonstrate that both the success and failure of SSMs in these domains can be explained by studying the role of the gating mechanism, a prevalent component in modern recurrent networks. Specifically, we show through theory and experiments that this gating mechanism causes SSMs to first learn an in-weights "memorization" solution, while delaying, or even preventing, convergence to a correct in-context learning solution. Importantly, this happens even in cases where there are no fundamental limitations due to the architecture or its memory capacity. On the other hand, we find that gating is often beneficial for improving generalization to long sequence lengths. Our results illuminate the crucial role of the gating mechanism in shaping both the training dynamics and generalization of SSMs, and provide a basis for understanding and improving linear-time models.
Sources
- Zoology: Measuring and Improving Recall in Efficient Language Models
- NVIDIA Nemotron Nano 2: An Accurate and Efficient Hybrid Mamba-Transformer Reasoning Model
- DeciMamba: Exploring the Length Extrapolation Potential of Mamba
- NVIDIA Nemotron 3: Efficient and Open Intelligence
- Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality
- Is Mamba Capable of In-Context Learning?
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces
- Efficiently Modeling Long Sequences with Structured State Spaces
- Understanding Input Selectivity in Mamba: Impact on Approximation Power, Memorization, and Associative Recall Capacity
- On the Recall Scaling Laws in Mamba: A Theoretical and Mechanistic Study via Hashing
- How Can Mamba Learn In Context with Outliers and Generalize Provably?
- Gating is Weighting: Understanding Gated Linear Attention through In-context Learning
- Decoupled Weight Decay Regularization
- Mamba Modulation: On the Length Generalization of Mamba
- To Infinity and Beyond: Tool-Use Unlocks Length Generalization in State Space Models
- Revisiting associative recall in modern recurrent models
- Can Mamba Learn How to Learn? A Comparative Study on In-Context Learning Tasks
- The mechanistic basis of data dependence and abrupt learning in an in-context classification task
- Understanding and Improving Length Generalization in Recurrent Models
- Mimetic Initialization Helps State Space Models Learn to Recall
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks