When Fancy Eviction Fails: Rethinking Cache Replacement For LLM Prefix Reuse
cs.DC, cs.LG
Submitted: 2026-09-24
Updated: 2026-09-24
Code: https://github.com/cacheMon/cache_
Terminology
Sources
- State of AI: An Empirical 100 Trillion Token Study with OpenRouter
- Scaling Laws for Neural Language Models
- Mixture-of-Experts Meets Instruction Tuning:A Winning Combination for Large Language Models
- Strata: Hierarchical Context Caching for Long Context Language Model Serving
Related papers
- iScheduler: Reinforcement Learning-Driven Continual Optimization for Large-Scale Resource Investment Problems
- SAMM: Sharded Automated Market Maker
- InferScale: GPU-Native KV Injection for Personalized LLM Serving
- Vigil: Accountable Liveness against Selective Silence
- Steelhead: Interleaving Partially Synchronous and Asynchronous Commit Rules on a Shared DAG
- Pushing CPU Speech Synthesis to the Wall: Extreme Inference Tuning under Serverless Architecture and Billing