Just Let Linear States Forget the Distant Past: Prefix Caching via Suffix Replay for Hybrid LLMs
cs.AI, cs.PF
Submitted: 2026-09-27
Updated: 2026-09-27
Code: https://github.com/vllm-project/vllm
Terminology
Sources
- Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality
- Cost-Efficient Large Language Model Serving for Multi-turn Conversations with CachedAttention
- Prompt Cache: Modular Attention Reuse for Low-Latency Inference
- MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework
- EPIC: Efficient Position-Independent Caching for Serving Large Language Models
- Kimi Linear: An Expressive, Efficient Attention Architecture
- Efficient Memory Management for Large Language Model Serving with PagedAttention
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
- Jamba: A Hybrid Transformer-Mamba Language Model
- LinearKV: One Cached State Suffices for Position-Independent Caching in Hybrid LLMs
- HYPIC: Accelerating Hybrid-Attention LLM Serving with Position-Independent Caching
- Olmo Hybrid: From Theory to Practice and Back
- MiniMax-01: Scaling Foundation Models with Lightning Attention
- Skeleton-of-Thought: Prompting LLMs for Efficient Parallel Generation
- Nemotron-H: A Family of Accurate and Efficient Hybrid Mamba-Transformer Models
- Training language models to follow instructions with human feedback
- Marconi: Prefix Caching for the Era of Hybrid LLMs
- Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
- Sparse Prefix Caching for Hybrid and Recurrent LLM Serving
- Preble: Efficient Distributed Prompt Scheduling for LLM Serving
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection