Sliding-window beats linear attention
cs.CL, cs.LG
Submitted: 2026-08-28
Updated: 2026-10-04
Code: https://github.com/tatsu-lab/stanford_alpaca
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Phi-4 Technical Report
- Phi-4-reasoning Technical Report
- gpt-oss-120b & gpt-oss-20b Model Card
- Longformer: The Long-Document Transformer
- Transformers to SSMs: Distilling Quadratic Knowledge to Subquadratic Models
- Llamba: Scaling Distilled Recurrent Models for Efficient Language Processing
- DiJiang: Efficient Large Language Models through Compact Kernelization
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
- Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality
- Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models
- RADLADS: Rapid Attention Distillation to Linear Attention Decoders at Scale
- The Llama 3 Herd of Models
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces
- LoRA: Low-Rank Adaptation of Large Language Models
- Mistral 7B
- HunyuanVideo: A Systematic Framework For Large Video Generative Models
- Liger: Linearizing Large Language Models to Gated Recurrent Structures
- Textbooks Are All You Need II: phi-1.5 technical report
- RWKV: Reinventing RNNs for the Transformer Era
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering