Scaling Parameter and Context in Attention: Native Sparse Attention from Mixture-of-Head
cs.CL
Submitted: 2026-09-30
Updated: 2026-09-30
Terminology
Sources
- Nemotron-4 340B Technical Report
- GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints
- LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding
- Sparse Delta Memory: Scaling the State of Linear RNNs through Sparsity
- Evaluating Large Language Models Trained on Code
- BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Training Verifiers to Solve Math Word Problems
- SwitchHead: Accelerating Transformers with Mixture-of-Experts Attention
- FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
- A Dataset of Information-Seeking Questions and Answers Anchored in Research Papers
- RoPE Distinguishes Neither Positions Nor Tokens in Long Contexts, Provably
- Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
- Attention Sink Forges Native MoE in Attention Layers: Sink-Aware Training to Address Head Collapse
- Transformer Feed-Forward Layers Are Key-Value Memories
- Active-Dormant Attention Heads: Mechanistically Demystifying Extreme-Token Phenomena in LLMs
- Measuring Massive Multitask Language Understanding
- Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps
- Training Compute-Optimal Large Language Models
- Mixtral of Experts
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering