CacheBridge: Efficient Cross-Model KV Cache Transfer
cs.AI
Submitted: 2026-09-01
Updated: 2026-09-01
Terminology
Sources
- Cross-Model KV Cache Transfer in LLM Families: A Closed-Form Linear Mapping for Prefill Reuse
- KV Cache Translation across Heterogeneous Large Language Models
- Efficient Multi-Adapter LLM Serving via Cross-Model KV-Cache Reuse with Activated LoRA
- SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer
- FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance
- RouteLLM: Learning to Route LLMs with Preference Data
- Latent Space Communication via K-V Cache Alignment
- Cache-to-Cache: Direct Semantic Communication Between Large Language Models
- PRISM: Fast Online LLM Serving via Scheduling-Memory Co-design
- Latent Cache Flow: Model-to-Model Communication Without Text
- Fathom-DeepResearch: Unlocking Long Horizon Information Retrieval and Synthesis for SLMs
- RoFormer: Enhanced Transformer with Rotary Position Embedding
- AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection