Tail-Replay: Escaping the Curse of Linear Attention in Prefix Caching for Hybrid LLMs
cs.LG, cs.AI
Submitted: 2026-08-31
Updated: 2026-08-31
Terminology
Sources
- LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding
- Cost-Efficient Large Language Model Serving for Multi-turn Conversations with CachedAttention
- Prompt Cache: Modular Attention Reuse for Low-Latency Inference
- MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework
- RULER: What's the Real Context Size of Your Long-Context Language Models?
- EPIC: Efficient Position-Independent Caching for Serving Large Language Models
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
- Jamba: A Hybrid Transformer-Mamba Language Model
- HYPIC: Accelerating Hybrid-Attention LLM Serving with Position-Independent Caching
- LinearKV: One Cached State Suffices for Position-Independent Caching in Hybrid LLMs
- Olmo Hybrid: From Theory to Practice and Back
- MiniMax-01: Scaling Foundation Models with Lightning Attention
- Nemotron-H: A Family of Accurate and Efficient Hybrid Mamba-Transformer Models
- Training language models to follow instructions with human feedback
- Marconi: Prefix Caching for the Era of Hybrid LLMs
- Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
- Samba: Simple Hybrid State Space Models for Efficient Unlimited Context Language Modeling
- Sparse Prefix Caching for Hybrid and Recurrent LLM Serving
- ProphetKV: User-Query-Driven Selective Recomputation for Efficient KV Cache Reuse in Retrieval-Augmented Generation
- Gated Delta Networks: Improving Mamba2 with Delta Rule
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks