KV-Lingo: Learning KV-Cache Translators with Distillation
cs.CL, cs.LG
Submitted: 2026-09-26
Updated: 2026-09-26
Terminology
Sources
- See What I See, Know What I Think: Dense Latent Communication Across Heterogeneous Agents
- QuAC : Question Answering in Context
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Training Verifiers to Solve Math Word Problems
- DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
- Latent Space Communication via K-V Cache Alignment
- Don't be lazy: CompleteP enables compute-efficient deep transformers
- Stochastic KV Routing: Enabling Adaptive Depth-Wise Cache Sharing
- Training Transformers for KV Cache Compressibility
- RelayCaching: Accelerating LLM Collaboration via Decoding KV Cache Reuse
- Cross-Model KV Cache Transfer in LLM Families: A Closed-Form Linear Mapping for Prefill Reuse
- Training Compute-Optimal Large Language Models
- KV Prediction for Improved Time to First Token
- RULER: What's the Real Context Size of Your Long-Context Language Models?
- Mistral 7B
- Kimi K2: Open Agentic Intelligence
- Kimi Linear: An Expressive, Efficient Attention Architecture
- A Universal Context-Reuse Layer for Cross-Model KV Sharing
- Convergent Learning: Do different neural networks learn the same representations?
- DroidSpeak: KV Cache Sharing for Cross-LLM Communication and Multi-LLM Serving
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering