KV-COBRA: KV Cache Compression via Co-Optimized Bit-Rank Allocation
cs.LG, cs.AI
Submitted: 2026-09-21
Updated: 2026-09-21
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- LLM Inference Unveiled: Survey and Roofline Model Insights
- Efficient Streaming Language Models with Attention Sinks
- GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
- SVDq: 1.25-bit and 410x Key Cache Compression for LLM Attention
- KQ-SVD: Compressing the KV Cache with Provable Guarantees on Attention Fidelity
- TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate
- Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models
- KVzip: Query-Agnostic KV Cache Compression with Context Reconstruction
- DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
- Dynamic Memory Compression: Retrofitting LLMs for Accelerated Inference
- Quantizing With Randomized Hadamard Transforms: Efficient Heuristic Now Proven
- The Llama 3 Herd of Models
- Mistral 7B
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- RULER: What's the Real Context Size of Your Long-Context Language Models?
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks