Nonequilibrium Phases of Repulsive Self-Attention: Chaos, Attention Condensation, and Emergent Locality
cond-mat.dis-nn, cond-mat.stat-mech, cs.LG
Submitted: 2026-09-23
Updated: 2026-09-23
Terminology
Sources
- Attention Is All You Need
- Universal Transformers
- Deep Equilibrium Models
- Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach
- The emergence of clusters in self-attention dynamics
- A mathematical perspective on Transformers
- Attention is Not All You Need: Pure Attention Loses Rank Doubly Exponentially with Depth
- Signal Propagation in Transformers: Theoretical Perspectives and the Role of Rank Collapse
- Clustered Attractor Manifolds and Dynamical Condensation in Self-Attention
- Recurrent Self-Attention Dynamics: An Energy-Agnostic Perspective from Jacobians
- Dynamical Mean-Field Theory of Self-Attention Neural Networks
- Efficient Streaming Language Models with Attention Sinks
- Massive Activations in Large Language Models
- Stabilizing Transformer Training by Preventing Attention Entropy Collapse
- In-context Learning and Induction Heads
- Light-cone spreading of perturbations and the butterfly effect in a classical spin chain
Related papers
- Few-Shot Neuromorphic Vision in a Nonlinear Photonic Network Laser
- Hyperbolic lattices with mass disorder: Phases and phase transitions
- Adaptive Neural Quantum States: A Recurrent Neural Network Perspective
- Machine learning Majorana topology using unsupervised and supervised learning
- Quenched fluctuation-induced force arising from polarization disorder