Where Activation Sparsity and KV-Cache Sparsity Cross in LLM Decoding
cs.LG, cs.CL, cs.PF
Submitted: 2026-09-27
Updated: 2026-09-27
Code: https://github.com/js-lee-AI/ByteCross
Terminology
Sources
- Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
- GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints
- Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads
- PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling
- KnapSpec: Self-Speculative Decoding via Adaptive Layer Selection as a Knapsack Problem
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
- LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding
- The Llama 3 Herd of Models
- FlashDecoding++: Faster Large Language Model Inference on GPUs
- Efficient Memory Management for Large Language Model Serving with PagedAttention
- CATS: Contextually-Aware Thresholding for Sparsity in Large Language Models
- DART: Draft-Agreement Routing for Training-Free Adaptive Thinking Budgets in Hybrid Reasoning Models
- Fast Inference from Transformers via Speculative Decoding
- SnapKV: LLM Knows What You are Looking for Before Generation
- EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty
- AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration
- Deja Vu: Contextual Sparsity for Efficient LLMs at Inference Time
- KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache
- Pointer Sentinel Mixture Models
- ReLU Strikes Back: Exploiting Activation Sparsity in Large Language Models
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks