Switching Linear Attention
cs.LG, cs.AI, cs.CL
Submitted: 2026-09-30
Updated: 2026-09-30
Code: https://github.com/fla-org/flash-linear-attention
Terminology
Sources
- In-Context Language Learning: Architectures and Algorithms
- Language Models Enable Simple Systems for Generating Structured Views of Heterogeneous Data Lakes
- Just read twice: closing the recall gap for recurrent language models
- Titans: Learning to Memorize at Test Time
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Scaling Context Requires Rethinking Attention
- A Unifying Framework for Parallelizing Sequential Models with Linear Dynamical Systems
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces
- Longhorn: State Space Models are Amortized Online Learners
- Why Are Linear RNNs More Parallelizable?
- M$^2$RNN: Non-Linear RNNs with Matrix-Valued States for Scalable Language Modeling
- Revisiting associative recall in modern recurrent models
- Gated KalmaNet: A Fading Memory Layer Through Test-Time Ridge Regression
- WinoGrande: An Adversarial Winograd Schema Challenge at Scale
- Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
- Learning State-Tracking from Code Using Linear RNNs
- End-to-End Test-Time Training for Long Context
- Kimi Linear: An Expressive, Efficient Attention Architecture
- LLaMA: Open and Efficient Foundation Language Models
- Uncovering mesa-optimization algorithms in Transformers
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks