Kascade: A Practical Sparse Attention Method for Long-Context LLM Inference
cs.LG, cs.AI, cs.DC
Submitted: 2025-12-18
Updated: 2026-08-31
Code: https://github.com/microsoft/kascade
Terminology
Sources
- gpt-oss-120b & gpt-oss-20b Model Card
- GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints
- Longformer: The Long-Document Transformer
- Leveraging redundancy in attention with Reuse Transformers
- SeerAttention: Learning Intrinsic Sparse Attention in Your LLMs
- SeerAttention-R: Sparse Attention Adaptation for Long Reasoning
- The Llama 3 Herd of Models
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- What Matters in Transformers? Not All Attention is Needed
- TurboRAG: Accelerating Retrieval-Augmented Generation with Precomputed KV Caches for Chunked Text
- Qwen3 Technical Report
- Less Is More: Fast and Accurate Reasoning with Cross-Head Unified Sparse Attention
- LServe: Efficient Long-sequence LLM Serving with Unified Sparse Attention
- Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference
- LazyFormer: Self Attention with Lazy Update
- Gemma: Open Models Based on Gemini Research and Technology
- TileLang: A Composable Tiled Programming Model for AI Systems
- Efficient Streaming Language Models with Attention Sinks
- Sharing Attention Weights for Fast Transformer
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks