PrefillShare: A Shared Prefill Module for KV Reuse in Multi-LLM Disaggregated Serving
cs.LG, cs.DC
Submitted: 2026-02-12
Updated: 2026-09-26
Code: https://github.com/langchain-ai/langchain
Terminology
Sources
- GPT-4 Technical Report
- SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
- FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance
- Evaluating Large Language Models Trained on Code
- Training Verifiers to Solve Math Word Problems
- FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
- LMFlow: An Extensible Toolkit for Finetuning and Inference of Large Foundation Models
- The Llama 3 Herd of Models
- The Cost of Dynamic Reasoning: Demystifying AI Agents and Test-Time Scaling from an AI Infrastructure Perspective
- GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers
- DroidSpeak: KV Cache Sharing for Cross-LLM Communication and Multi-LLM Serving
- LMCache: An Efficient KV Cache Layer for Enterprise-Scale LLM Inference
- Decoupled Weight Decay Regularization
- WizardCoder: Empowering Code Large Language Models with Evol-Instruct
- KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows
- Small LLMs Are Weak Tool Learners: A Multi-LLM Agent
- ToolOrchestra: Elevating Intelligence via Efficient Model and Tool Orchestration
- Gemini: A Family of Highly Capable Multimodal Models
- Mixture-of-Agents Enhances Large Language Model Capabilities
- Plan-and-Solve Prompting: Improving Zero-Shot Chain-of-Thought Reasoning by Large Language Models
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks