The Attention Within: Consensus Dynamics in Selective State Space Models
cs.LG, cs.AI, cs.SY, eess.SY, math.DS, math.OC
Submitted: 2026-09-16
Updated: 2026-09-16
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: Selective state space models (SSMs) have recently emerged as a compelling alternative to transformers, combining competitive performance with substantially improved inference efficiency.
Terminology
Abstract
Selective state space models (SSMs) have recently emerged as a compelling alternative to transformers, combining competitive performance with substantially improved inference efficiency. At each SSM layer, a sequence of hidden states are propagated by a recurrence, mixing information of different tokens. Despite using a different mechanism, this mixing plays a role analogous to attention in transformers. In fact, recent works have shown that the two architectures may be closer than they first appear, as this recurrence admits a formulation akin to linear attention. In transformers, attention is known to drive the tokens to cluster, i.e., to reach consensus, collapsing in the limit to a single direction. Thus, we ask: does the recurrence at the core of SSMs drive the tokens to consensus, as attention does in transformers? To answer this question, we take a dynamical systems perspective on SSMs, modeling the evolution of tokens across layers as an ordinary differential equation. By exploiting input-to-state stability arguments, we establish local exponential stability of the consensus equilibria and characterize their domain of attraction for time-varying weight matrices, a setting not addressed by previous results. We thereby show that the resemblance between SSMs and transformers does run deeper: the recurrence at the core of SSMs aggregates tokens just as attention does. Numerical experiments on a pretrained Mamba-2 model point to the output gate as the component that regulates the extent of this consensus, preventing the tokens from reaching it in full.
Sources
- GPT-4 Technical Report
- Neural Machine Translation by Jointly Learning to Align and Translate
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- An Empirical Study of Mamba-based Language Models
- The Llama 3 Herd of Models
- Clustering in pure-attention hardmax transformers and its role in sentiment analysis
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces
- The Illusion of State in State-Space Models
- A Comparative Analysis of Contextual Representation Flow in State-Space and Transformer Architectures
- The Asymptotic Behavior of Attention in Transformers
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks