On-Demand Attention: Language Models Know When to Recall
cs.CL
Submitted: 2026-09-17
Updated: 2026-09-26
Comments: 28 pages, 5 figures
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- SpotAttention: Plug-In Block-Sparse Routing for Pretrained Long-Context Transformers
- Learning When to Attend: Conditional Memory Access for Long-Context LLMs
- Rag Performance Prediction for Question Answering
- SeerAttention-R: Sparse Attention Adaptation for Long Reasoning
- Gemma 4 Technical Report
- Language Models Can Control Their Own Attention
- Self-Selected Attention Span for Accelerating Large Language Model Inference
- Efficient Memory Management for Large Language Model Serving with PagedAttention
- Decoupled Weight Decay Regularization
- Learning When Not to Attend Globally
- LoGo: Token-Level Dynamic Local-Global Attention
- YaRN: Efficient Context Window Extension of Large Language Models
- Qwen3 Technical Report
- GLU Variants Improve Transformer
- NLL-Guided Full-Attention Layer Selection for Training-Free Sliding-Window Adaptation
- Short-Context Dominance: How Much Local Context Natural Language Actually Needs?
- Attention Is All You Need
- Switch Attention: Towards Dynamic and Fine-grained Hybrid Transformers
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering