The Evolution of Attention in Large Language Models: Mechanisms, Trade-offs, and Emerging Trends
cs.CL
Submitted: 2026-09-30
Updated: 2026-09-30
Code: https://github.com/meta-llama/llama-models
Terminology
Sources
- Attention Is All You Need
- Language Models are Few-Shot Learners
- LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding
- RULER: What's the Real Context Size of Your Long-Context Language Models?
- DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
- Kwai Summary Attention Technical Report
- Fast Transformer Decoding: One Write-Head is All You Need
- GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints
- Reducing Transformer Key-Value Cache Size with Cross-Layer Attention
- Longformer: The Long-Document Transformer
- Big Bird: Transformers for Longer Sequences
- Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention
- On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability
- Hierarchical Sparse Attention Done Right: Toward Infinite Context Modeling
- IndexCache: Accelerating Sparse Attention via Cross-Layer Index Reuse
- LongCat Sparse Attention: Taming the Lightning via Streaming-aware Hierarchical Cross-Layer Indexing
- Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention
- Retentive Network: A Successor to Transformer for Large Language Models
- Linear Transformers Are Secretly Fast Weight Programmers
- Gated Delta Networks: Improving Mamba2 with Delta Rule
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering