Opinion Leader Dynamics: How Sparse Attention Shapes Token Clustering
cs.LG, cs.AI
Submitted: 2026-09-21
Updated: 2026-09-21
Comments: Code is available at https://github.com/Jingkun-Liu/Opinion-Leader-Dynamics.git
Code: https://github.com/Jingkun-Liu/Opinion-Leader-Dynamics
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: Sparse attention reduces the quadratic cost of global self-attention while retaining strong empirical performance, but how its restricted interactions shape the evolution of token representations
Terminology
Abstract
Sparse attention reduces the quadratic cost of global self-attention while retaining strong empirical performance, but how its restricted interactions shape the evolution of token representations remains theoretically underexplored. Modeling tokens as particles on the unit sphere, we introduce opinion leader dynamics, a framework that identifies two mechanisms through which token groups converge internally while maintaining distinct limiting directions. In the explicit model, fixed representatives induce a potential that attracts tokens toward distinct local maxima. In the implicit model, disconnected interaction groups evolve toward separate consensus directions. We formulate both models as reverse Wasserstein gradient flows and establish exponential convergence under suitable conditions. We further connect these theoretical predictions to token evolution in frontier sparse-attention LLMs that motivate our framework. Across four benchmarks, Kimi-K3, MiniMax-M3, and DeepSeek-V4-Flash consistently exhibit clearer cluster separation and higher clustering scores than the dense-attention model GLM-4.7-Flash in projected token representations. These observations support the relevance of the predicted multiple-group structure to trained frontier LLMs, while finite-particle simulations illustrate the theoretical convergence behavior. Together, our results connect restricted token interactions to distinct group-level attractors, providing a dynamical account of how sparse attention can support alignment within groups while preserving separation between them.
Sources
- GPT-4 Technical Report
- Longformer: The Long-Document Transformer
- Evaluating Large Language Models Trained on Code
- Quantitative Clustering in Mean-Field Transformer Models
- Generating Long Sequences with Sparse Transformers
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- A mathematical perspective on Transformers
- Measuring Mathematical Problem Solving With the MATH Dataset
- RULER: What's the Real Context Size of Your Long-Context Language Models?
- Clustering in Causal Attention Masking
- Reformer: The Efficient Transformer
- FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference
- MiniMax Sparse Attention
- Understanding and Improving Length Generalization in Hierarchical Sparse Attention Models
- Characterizing the Expressivity of Local Attention in Transformers
- DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
- Krause Synchronization Transformers
- MoBA: Mixture of Block Attention for Long-Context LLMs
- UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction
- Landmark Attention: Random-Access Infinite Context Length for Transformers
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks