Evaluating the accuracy of KV cache reuse techniques
cs.LG
Submitted: 2026-09-25
Updated: 2026-09-25
Code: https://github.com/boxoffice1280/boxoffice1280
Terminology
Sources
- Sparse Attention across Multiple-context KV Cache
- An experimental study of KV cache reuse strategies in chunk-level caching systems
- Benchmarking Large Language Models in Retrieval-Augmented Generation
- Benchmarking Deep Search over Heterogeneous Enterprise Data
- WixQA: A Multi-Dataset Benchmark for Enterprise Retrieval-Augmented Generation
- Cartridges: Lightweight and general-purpose long context representations via self-study
- EVICPRESS: Joint KV-Cache Compression and Eviction for Efficient LLM Serving
- Characterising Bias in Compressed Models
- EPIC: Efficient Position-Independent Caching for Serving Large Language Models
- Mistral 7B
- LMCache: An Efficient KV Cache Layer for Enterprise-Scale LLM Inference
- The Llama 3 Herd of Models
- TurboRAG: Accelerating Retrieval-Augmented Generation with Precomputed KV Caches for Chunked Text
- Large Language Models Sensitivity to The Order of Options in Multiple-Choice Questions
- NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark
- Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
- From Prefix Cache to Fusion RAG Cache: Accelerating LLM Inference in Retrieval-Augmented Generation
- ProphetKV: User-Query-Driven Selective Recomputation for Efficient KV Cache Reuse in Retrieval-Augmented Generation
- RAGPulse: An Open-Source RAG Workload Trace to Optimize RAG Serving Systems
- Simple synthetic data reduces sycophancy in large language models
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks