Draft-KV: Learning Useful Latent Communication Between Language Models
cs.CL
Submitted: 2026-09-28
Updated: 2026-09-28
Code: https://github.com/Svardfox/Draft-KV
Terminology
Sources
- See What I See, Know What I Think: Dense Latent Communication Across Heterogeneous Agents
- Adapting Language Models to Compress Contexts
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Improving Factuality and Reasoning in Language Models through Multiagent Debate
- Enabling Agents to Communicate Entirely in Latent Space
- Mixture of Thoughts: Learning to Aggregate What Experts Think, Not Just What They Say
- Cache-to-Cache: Direct Semantic Communication Between Large Language Models
- In-context Autoencoder for Context Compression in a Large Language Model
- Large Language Model based Multi-Agents: A Survey of Progress and Challenges
- Training Large Language Models to Reason in a Continuous Latent Space
- Measuring Massive Multitask Language Understanding
- LoRA: Low-Rank Adaptation of Large Language Models
- Agent Primitives: Reusable Latent Building Blocks for Multi-Agent Systems
- When Less Latent Leads to Better Relay: Information-Preserving Compression for Latent Multi-Agent LLM Collaboration
- DroidSpeak: KV Cache Sharing for Cross-LLM Communication and Multi-LLM Serving
- Learning to Compress Prompts with Gist Tokens
- Let Models Speak Ciphers: Multiagent Debate through Embeddings
- Qwen2.5 Technical Report
- Communicating Activations Between Language Model Agents
- KVComm: Enabling Efficient LLM Communication through Selective KV Sharing
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering