DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression
cs.CL
Submitted: 2026-09-17
Updated: 2026-09-17
Code: https://github.com/anomalyco/opencode
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints
- IndexCache: Accelerating Sparse Attention via Cross-Layer Index Reuse
- PowerAttention: Exponentially Scaling of Receptive Fields for Effective Sparse Attention
- BabyVision: Visual Reasoning Beyond Language
- Evaluating Large Language Models Trained on Code
- DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation
- Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models
- Training Verifiers to Solve Math Word Problems
- DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models
- DeepSeek-V3 Technical Report
- DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
- DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
- DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
- RMNP: Row-Momentum Normalized Preconditioning for Scalable Matrix-Based Optimization
- NL2Repo-Bench: Towards Long-Horizon Repository Generation Evaluation of Coding Agents
- SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines
- HySparse: A Hybrid Sparse Attention Architecture with Oracle Token Selection and KV Cache Sharing
- Chartography: A Benchmark for Professional Chart Understanding
- Memory-Efficient LLM Pretraining via Minimalist Optimizer Design
- SimpleQA Verified: A Reliable Factuality Benchmark to Measure Parametric Knowledge
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering