ValueDiff: Value-Geometric KV Cache Eviction for Sink-Suppressed LLMs
cs.LG, cs.AI
Submitted: 2026-09-20
Updated: 2026-09-20
Comments: 10 pages, 3 Figues
Code: https://github.com/gkamradt/LLMTest_NeedleInAHaystack
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Apple Intelligence Foundation Language Models
- Fast On-device LLM Inference with NPUs
- Accelerating Mobile Language Model via Speculative Decoding and NPU-Coordinated Execution
- Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation
- Why do LLMs attend to the first token?
- ManifoldKV: Training-Free KV Cache Compression via Euclidean Outlier Detection
- A Simple and Effective $L_2$ Norm-Based Strategy for KV Cache Compression
- Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
- gpt-oss-120b & gpt-oss-20b Model Card
- Gemma 2: Improving Open Language Models at a Practical Size
- KVzip: Query-Agnostic KV Cache Compression with Context Reconstruction
- CAOTE: KV Cache Selection for LLMs via Attention Output Error-Based Token Eviction
- What are you sinking? A geometric approach on attention sink
- The Llama 3 Herd of Models
- Qwen2.5 Technical Report
- RULER: What's the Real Context Size of Your Long-Context Language Models?
- Measuring Mathematical Problem Solving With the MATH Dataset
- Let's Verify Step by Step
- Anisotropy Is Inherent to Self-Attention in Transformers
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks