DASC: Decay-Aware State Compression for Hybrid Linear-Attention Serving
cs.LG, cs.AI
Submitted: 2026-08-31
Updated: 2026-08-31
Code: https://github.com/MoonshotAI/Kimi-K3
Terminology
Sources
- Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces
- RULER: What's the Real Context Size of Your Long-Context Language Models?
- RAGCache: Efficient Knowledge Caching for Retrieval-Augmented Generation
- Jamba: A Hybrid Transformer-Mamba Language Model
- GPQA: A Graduate-Level Google-Proof Q&A Benchmark
- Samba: Simple Hybrid State Space Models for Efficient Unlimited Context Language Modeling
- MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
- Gated Delta Networks: Improving Mamba2 with Delta Rule
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks