Attention by Synchronization in Coupled Oscillator Networks
cs.LG, cs.NE, nlin.AO
Submitted: 2026-06-10
Updated: 2026-09-09
Journal ref: Transactions on Machine Learning Research, 2026
License: http://creativecommons.org/licenses/by/4.0/
The gist: We address transformer attention on energy-constrained physical substrates.
Terminology
Abstract
We address transformer attention on energy-constrained physical substrates. Softmax attention requires exponentiation and global reduction, operations with high energy cost on von Neumann hardware and no natural physical analog. We show that Kuramoto synchronization dynamics (which arise in electrical, mechanical, superconducting, and charge-density-wave oscillator arrays, among other physical systems) implement a well-defined attention operation. The resulting mechanism, fixed-query oscillator attention, replaces softmax's arithmetic with the equilibration of a gradient flow on the sphere: queries are learned anchors fixed on the sphere, and free oscillators evolve under Kuramoto--Lohe dynamics until they settle at positions encoding attention weights via cosine similarity. Because the computation is equilibration, no global exponential normalization is needed. The fixed point is provably unique and globally attractive from almost every initial condition, a guarantee that holds across every physical realization. Empirically, at the minimal hardware configuration (oscillator dimension d osc = 2), oscillator attention matches softmax on keyword spotting, and on subject-verb agreement it trains more reliably while reaching softmax's accuracy. Softmax retains an advantage on causal language modeling, but the gap decays as a power law in d osc. The main objective of this work is not to replace softmax in software but to provide a mathematically grounded blueprint for accurate attention on physical substrates.
Sources
- Speech Commands: A Dataset for Limited-Vocabulary Speech Recognition
- TinyStories: How Small Can Language Models Be and Still Speak Coherent English?
- Linformer: Self-Attention with Linear Complexity
- Generating Long Sequences with Sparse Transformers
- Longformer: The Long-Document Transformer
- SpikeGPT: Generative Pre-trained Language Model with Spiking Neural Networks
- SpikeBERT: A Language Spikformer Learned from BERT with Knowledge Distillation
- Emergence Transformer: Dynamical Temporal Attention Matters
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks