LatentPort: Beyond KV Cache - Cross-Model Transfer of Recurrent Memory in Hybrid Language Models: A 4B-to-9B Hybrid-State Handoff Without Target Prefix Replay
cs.CL, cs.AI
Submitted: 2026-09-07
Updated: 2026-09-07
Comments: 14 pages, 6 figures, 12 tables
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Cross-Model KV Cache Transfer in LLM Families: A Closed-Form Linear Mapping for Prefill Reuse
- KV Cache Translation across Heterogeneous Large Language Models
- Frozen Memory Is Not Enough: Rethinking External Memory as Extraction
- A Universal Context-Reuse Layer for Cross-Model KV Sharing
- HYPIC: Accelerating Hybrid-Attention LLM Serving with Position-Independent Caching
- Tail-Replay: Escaping the Curse of Linear Attention in Prefix Caching for Hybrid LLMs
- Semantic Cache Distillation: Efficient State Transfer via Reuse and Selective Patching
- Marconi: Prefix Caching for the Era of Hybrid LLMs
- The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale
- StateBridge: Training-free Hidden-state Alignment for Latent Communication in LLM Multi-Agent Systems
- CacheBridge: Efficient Cross-Model KV Cache Transfer
- Compressive Transformers for Long-Range Sequence Modelling
- Latent Cache Flow: Model-to-Model Communication Without Text
- Gated Delta Networks: Improving Mamba2 with Delta Rule
- S0 Tuning: Zero-Overhead Adaptation of Hybrid Recurrent-Attention Models
- WriteSAE: Sparse Autoencoders for Recurrent State
- DASC: Decay-Aware State Compression for Hybrid Linear-Attention Serving
- Metis: Memory Foundation Model
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering