Closing the Operational Gap in Semantic Caching
cs.IR, cs.CL, cs.LG
Submitted: 2026-06-18
Updated: 2026-09-01
Comments: 24 pages, 2 figures. Source code: https://github.com/aditeyabaral/operational-gap-semantic-caching. Models and Datasets: https://huggingface.co/redis. Accepted at EMNLP 2026, Industry Track
Code: https://github.com/aditeyabaral/operational-gap-semantic-caching
License: http://creativecommons.org/licenses/by/4.0/
The gist: Semantic caching cuts LLM inference costs by serving a cached response to semantically similar queries.
Terminology
Abstract
Semantic caching cuts LLM inference costs by serving a cached response to semantically similar queries. Standard practice evaluates these systems using PR-AUC, a metric that only measures how well scores rank and ignores whether they are usable at a fixed threshold. We show this mismatch leads to systematically poor deployment choices, as models with the highest PR-AUC are often the worst in operation. We introduce Precision--Cache Hit Ratio (P-CHR) AUC, a cache-aware metric that measures precision across cache utilization levels, and Operational Retention Rate (ORR), which captures how much offline ranking quality survives at deployment. We decompose the operational gap between offline and deployed quality into a recoverable threshold-utility component and an irreducible structural component fixed by the dataset's positive rate. Our experiments show that the threshold-utility gap is governed by the training objective rather than data scale, and yields only to re-normalizing scores over the candidate pool or changing the training objective. Ultimately, model selection for semantic caching is a threshold-utility problem, not a ranking one, and measuring it is the first step to closing the gap.
Sources
- From Exact Hits to Close Enough: Semantic Caching for LLM Embeddings
- ColBERT-Zero: To Pre-train Or Not To Pre-train ColBERT models
- GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
- Advancing Semantic Caching for LLMs with Domain-Specific Embeddings and Synthetic Data
- Towards General Text Embeddings with Multi-stage Contrastive Learning
- Semantic Caching for Low-Cost LLM Serving: From Offline Learning to Online Adaptation
- Jina Embeddings 2: 8192-Token General-Purpose Text Embeddings for Long Documents
- Passage Re-ranking with BERT
- Category-Aware Semantic Caching for Heterogeneous LLM Workloads
- Text Embeddings by Weakly-Supervised Contrastive Pre-training
- GPT Semantic Cache: Reducing LLM Costs and Latency via Semantic Embedding Caching
Related papers
- The Price of Isolation: Estimating the Ecosystem Cost of Symmetric Two-Sided A/B Testing
- SCAR: Semantic Continuity-Aware Retrieval for Efficient Context Expansion in RAG
- MixLoRA-DSI: Dynamically Expandable Mixture-of-LoRA Experts for Rehearsal-Free Generative Retrieval over Dynamic Corpora
- RRCM: Ranking-Driven Retrieval over Collaborative and Meta Memories for LLM Recommendation
- Right Family, Wrong Skill: Evaluating Risk Exposure in Agent Skill Retrieval
- UltRAG: a Universal Simple Scalable Recipe for Knowledge Graph RAG