Prefill-Free Cross-Family KV Cache Transfer for Heterogeneous Multi-Agent LLMs
cs.AI, cs.LG, cs.MA
Submitted: 2026-09-26
Updated: 2026-09-26
Code: https://github.com/huggingface/open-r1
Terminology
Sources
- Multi-Agent Consensus Seeking via Large Language Models
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- The Llama 3 Herd of Models
- Cross-Model KV Cache Transfer in LLM Families: A Closed-Form Linear Mapping for Prefill Reuse
- Robust Multi-Agent LLMs under Byzantine Faults
- Evaluating Very Long-Term Conversational Memory of LLM Agents
- RoFormer: Enhanced Transformer with Rotary Position Embedding
- Mixture-of-Agents Enhances Large Language Model Capabilities
- PrefillShare: A Shared Prefill Module for KV Reuse in Multi-LLM Disaggregated Serving
- Qwen3 Technical Report
- Recursive Multi-Agent Systems
- X-MAS: Towards Building Multi-Agent Systems with Heterogeneous LLMs
- Output-Aware Rotation for INT2 KV-Cache Quantization
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection