LongCat Sparse Attention: Taming the Lightning via Streaming-aware Hierarchical Cross-Layer Indexing
Wen Zan, Jiaqi Zhang, Jianchao Tan, Hong Liu, Cunguang Wang, Xiang Li, Duyue Ma, Guanyu Wu, Yifan Lu, Fengcun Li, Yerui Sun, Peng Pei, Yuchen Xie, Xunliang Cai
cs.AI, cs.CL, cs.DC, cs.LG
Submitted: 2026-08-03
Code: https://github.com/deepseek-ai/DeepSeek-V3.2-Exp
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
- GLM-5: from Vibe Coding to Agentic Engineering
- DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads
- TidalDecode: Fast and Accurate LLM Decoding with Position Persistent Sparse Attention
- Kascade: A Practical Sparse Attention Method for Long-Context LLM Inference
- HySparse: A Hybrid Sparse Attention Architecture with Oracle Token Selection and KV Cache Sharing
- Scaling Embeddings Outperforms Scaling Experts in Language Models
- Generating Long Sequences with Sparse Transformers
- Longformer: The Long-Document Transformer
- RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval
- MoBA: Mixture of Block Attention for Long-Context LLMs
- Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention
- Fast Transformer Decoding: One Write-Head is All You Need
- DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
- DeepSeek-V3 Technical Report
- LongCat-Flash Technical Report
- IndexCache: Accelerating Sparse Attention via Cross-Layer Index Reuse
- HELMET: How to Evaluate Long-Context Language Models Effectively and Thoroughly
- RULER: What's the Real Context Size of Your Long-Context Language Models?
- MS MARCO: A Human Generated MAchine Reading COmprehension Dataset
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection