InferScale: GPU-Native KV Injection for Personalized LLM Serving
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "InferScale: GPU-Native KV Injection for Personalized LLM Serving".
Jane: Large language models are increasingly deployed with persistent personalized context, such as accumulated memory profiles or long conversation histories, that is shared across a user’s many requests.
Tom: First, who's behind it and why it matters.
Title and authors: Tom: So, let's talk about the title and who wrote this paper, InferScale: GPU-Native KV Injection for Personalized LLM Serving. Basically, the title tells us right away that they are focusing on putting Key-Value states directly into the GPU's cache instead of having to repeatedly prefill the prompt text.
Jane: That makes perfect sense when you think about how much time gets spent redoing that same initial setup for every single request; it sounds like they’re trying to eliminate that bottleneck entirely.
Lu: The authors, Peter Li and Prashant Pandey, are from Northeastern University, and their work here really bridges the gap between data systems design and direct LLM inference architecture.
Meng: I wonder how much of the complexity is baked into making it "GPU-Native" versus just being a clever software trick running on top of existing frameworks.
Lalam: For me, the fact that they are focusing on KV injection suggests a more fundamental architectural improvement in how we manage state during inference, which could lead to far more efficient and consistent personalized AI experiences for everyone.
The paper's summary: Tom: So, to summarize what InferScale actually does, the authors present a GPU-native LLM memory system that completely swaps out the old way of repeatedly prefilling prompt text with a reusable KV state. They do this by precomputing each memory fact’s Key-Value representation and injecting it directly into vLLM’s paged cache.
Jane: So, instead of the engine having to re-read and re-process that stored context every time, it's already there, ready to be used instantly when a relevant piece of memory is retrieved. That’s a significant efficiency gain for handling long conversation histories or large RAG documents.
Lu: The paper details this through two main phases: an offline phase where they build the reusable KV state, and an online serving phase that handles semantic retrieval and then injects those facts into the model's attention layer.
Meng: I see their methodology involves a clever separation between storing the semantic embedding for retrieval and storing the pre-RoPE KV representation for injection, which they link together using a shared fact id. That sounds like some pretty intricate data management to get right on hardware.
Lalam: The concept of separating the search space—the semantic embedding space for finding facts and the KV space for actual conditioning—is brilliant because it allows retrieval to happen in one area while conditioning happens in another, which is a very efficient way to handle personalization at scale.
The paper's improvements: Tom: Now, let’s look at the specific improvements they highlight. They introduce Chunked RoPE for position-independent KV reuse and Context-Window Encoding to make sure the retrieved facts are actually accurate and usable in context.
Jane: Chunked RoPE is interesting because it lets them store keys before rotary position encoding and then re-rotate them at insertion time so that they behave exactly like they were prefilled at their original spot, no matter where you inject them.
Lu: And the Context-Window Encoding is a necessary fix; it addresses the accuracy loss that happens when you encode each fact completely independently by bundling it with a small window of preceding conversation turns.
Meng: From an engineering viewpoint, the challenge they solved with Chunked RoPE is making sure that relocating those precomputed keys doesn't mess up the attention scores at all, which is a tricky thing to get right when dealing with positional encodings.
Lalam: The Context-Window Encoding ensures that even though we are retrieving facts separately, we aren't losing the surrounding conversational context needed for disambiguation; this makes the retrieved memory much more reliable for use.
Conclusion: Tom: So, wrapping things up on this InferScale paper, it really boils down to proving that retrieved memory doesn't need to be re-prefilled every time; its KV representation can be injected directly into the attention cache, which is an exact equivalence to prompt injection for any positional encoding scheme.
Jane: That means we get near-constant time-to-first-token latency regardless of how large the retrieved context becomes, which is a huge win for real applications where context size varies wildly.
Lu: The implications are quite broad; this architecture supports dynamically assembling complex, non-contiguous context blocks from retrieved facts and injecting them precisely where they are needed in the prompt.
Meng: Practically speaking, keeping all the retrieval indices and KV stores on the GPU minimizes those slow host-to-device PCIe transfers, which is essential for handling high concurrency without massive latency spikes.
Lalam: I think this architecture fundamentally enhances how we design personalized AI; it moves us toward systems where memory management isn't a performance constraint but an integrated part of the inference process.
Tom: What we've seen here with InferScale is that sophisticated memory handling can lead to much more responsive and scalable personalized AI serving. We’re really looking forward to seeing how this kind of direct KV injection plays out in production environments.
Northeastern University
cs.DC, cs.LG
Submitted: 2026-07-29
Updated: 2026-09-27
Code: https://github.com/saltsystemslab/InferScale
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 92/100
The gist: Large language models are increasingly deployed with persistent personalized context, such as accumulated memory profiles or long conversation histories, that is shared across a user’s many requests.
Key concepts
- Context-Window Encoding
- This technique encodes each stored memory fact along with a small window of previous conversation turns. This ensures that when a fact is retrieved, it retains enough local context to be accurate without needing to recompute the entire prompt or fine-tune the model.
- Chunked RoPE
- This method stores keys before rotary position encoding and then re-rotates them at insertion time during serving. This allows a precomputed fact's KV representation to be placed at any position in the sequence while ensuring it produces identical attention scores as if it had been prefilled there originally.
- KV Injection via vLLM KV Connector
- This is the mechanism that places the precomputed memory directly into vLLM's paged cache. It bypasses traditional prefill steps by registering the memory during scheduling and performing a direct GPU-to-GPU copy, ensuring tokens are ready immediately.
Terminology
Summary
Large language models are increasingly deployed with persistent personalized context, such as accumulated memory profiles or long conversation histories, that is shared across a user’s many requests. InferScale presents a GPU-native LLM memory system that replaces repeated prompt prefilling with reusable KV state by precomputing each memory fact’s KV representation and injecting it directly into vLLM’s paged cache, leading to near-constant time-to-first-token (TTFT) regardless of the retrieved context size.
The gist
InferScale is a GPU-native LLM memory system that replaces repeated prompt prefilling with reusable KV state by precomputing each memory fact’s KV representation and injecting it directly into vLLM’s paged cache, leading to near-constant time-to-first-token (TTFT) regardless of the retrieved context size.
How it works
InferScale is designed as a data-systems co-design
that separates memory processing into two phases: an offline phase and an online serving phase. The offline phase involves constructing reusable KV state. This includes:
-
Context-Window Encoding: Each fact is encoded together with a configurable window of preceding conversation turns, while only the KV corresponding to the fact is retained. This preserves local contextual information while allowing every fact to remain independently retrievable and reusable, recovering nearly the accuracy of prompt injection without serving-time recomputation or model fine-tuning.
-
Persistent Memory Store: Each fact is represented in two complementary forms linked by a shared fact id: a semantic embedding optimized for approximate nearest-neighbor search, and a reusable KV representation that conditions the LLM during inference. This separation allows retrieval to run in an embedding space while conditioning runs in the model’s KV space.
How it works (Continued)
The online serving phase executes for every request and performs three key steps:
-
Semantic Retrieval: The incoming query is embedded, and a top-k search is performed using Jasper, a GPU-resident graph-based Approximate Nearest Neighbor (ANN) index built over semantic embeddings. This retrieves the relevant facts identified by fact ids.
-
Chunked RoPE Composition: Retrieved facts are resolved to their pre-RoPE KV representations, prepended with a fixed instruction prefix, and composed into a single memory segment. This composition assigns contiguous virtual positions (0 to m-1) and applies RoPE to the stored keys on the fly, ensuring that relocated keys yield exactly the same attention scores as if they had been prefilled at their original position.
-
KV Injection via vLLM KV Connector: The composed memory is injected directly into vLLM’s paged cache using a specialized connector plugin. This involves registering the composed memory during scheduling so that paged blocks are allocated without scheduling them for prefill, and then performing a GPU-to-GPU copy of the composed KV directly into those blocks before the query tokens are prefilled.
Key Technical Innovations
The system addresses specific challenges inherent in memory reuse:
-
Chunked RoPE: This technique stores keys before rotary position encoding and re-rotates them at insertion time, allowing a fact encoded once to be injected at any prompt position while preserving exactly the same attention behavior as prompt injection.
-
Context-Window Encoding: This mitigates the accuracy loss from encoding facts independently by encoding each fact together with a small window of preceding conversation turns, which recovers nearly the accuracy of prompt injection without requiring serving-time recomputation or cache re-encoding.
Performance and Evaluation
InferScale demonstrates significant performance improvements over existing systems like Mem0:
(RQ1) Serving Latency:
(RQ2) Accuracy:
Across three open-weight models on LoCoMo, InferScale keeps engine TTFT nearly constant as the retrieval budget grows (3.6–4.8× lower than Mem0 at k=50), and raises throughput by 3.7–4.5× at 100 concurrent users, while context-window encoding recovers accuracy to within a few points of prompt injection (60.3% vs. 63.3%).
(RQ4) CPU Offloading:
Offloading the pre-computed KV embeddings to host DRAM further lifts the GDDR capacity bound at only a few-millisecond TTFT cost, with end-to-end latencies growing by tens of milliseconds as retrieval budget increases.
Contributions
The contributions are:
-
Retrieve-and-inject inference: Proving that retrieved memory need not be re-prefilled on every request; its KV representation can be injected directly into the attention cache, and proving this is an exact equivalence to prompt injection at query positions for any positional encoding scheme (Theorem 3).
-
Chunked RoPE for position-independent KV reuse: Introducing a method that allows stored keys to be relocated exactly without breaking attention scores.
Improvements for AI systems
Here are specific improvements to AI systems based on the InferScale paper, detailing what these improved systems can achieve:
)1. Transition from Prompt-Injection Memory Systems to KV-Injection Memory Systems:
By replacing traditional prompt injection (where retrieved memory is serialized and re-prefilled as prompt text) with InferScale's KV injection mechanism, the system achieves a fundamental decoupling of memory retrieval cost from context size.
-
An improved AI system can handle significantly larger personalized contexts (e.g., long conversation histories or large RAG documents) without incurring a quadratic increase in Time-to-First-Token (TTFT). This is critical for long-running, stateful conversational agents or complex document analysis where the
memory
is extensive. -
The system can maintain near-constant TTFT even when the retrieval budget increases from small to large, as the latency scales linearly with query length rather than quadratically with memory size.
)2. Enable GPU-Native, High-Throughput Personalized Serving:
InferScale is designed as a GPU-native engine integrated via vLLM's KV connector interface, keeping all retrieval indices and KV stores on the GPU (GDDR/HBM).
-
The improved system can support massive concurrent user loads (e.g., 100+ users) with near-linear scaling of throughput, as the high cost of prompt re-filling is eliminated.
-
By keeping data resident on the GPU, it minimizes costly host-to-device PCIe transfers during retrieval/injection, leading to a substantial reduction in end-to-end query latency compared to CPU/host memory solutions.
)3. Achieve Exact Positional Relocation for Dynamic Context:
The introduction of Chunked RoPE allows retrieved memory facts (which may appear at arbitrary prompt positions) to be injected exactly where they are needed, preserving the original attention behavior.
- The AI can dynamically assemble complex, non-contiguous context blocks from retrieved facts and inject them into the prompt without requiring model fine-tuning or expensive re-encoding. This allows for highly flexible memory assembly tailored precisely to the immediate query's requirements.
)4. Enhance Contextual Accuracy via Context-Window Encoding:
The use of Context-Window Encoding ensures that retrieved facts are not encoded in isolation, but rather with a small window of preceding conversation turns, preserving necessary contextual disambiguation information.
- This allows the system to maintain high factual recall accuracy (approaching prompt injection levels) even when facts are retrieved independently. This is crucial for resolving ambiguous references (e.g., pronouns or deictic terms) within a long memory span, ensuring that the injected KV state is semantically rich and accurate.
)5. Optimize Serving Latency via Hybrid Storage Strategies:
The paper demonstrates that offloading the pre-RoPE KV store to host DRAM over PCIe provides a practical way to scale capacity beyond GPU HBM limits while keeping TTFT relatively low (a few milliseconds).
- The improved system can support much larger total memory stores than physically fit in GPU VRAM by intelligently spilling
cold
or less immediately relevant KV states to host DRAM. This allows the system to serve a much larger personalized knowledge base than would otherwise be possible, at a manageable cost that scales with retrieval budget.
)6. Create an Evolving, Incrementally Updatable Memory System:
Since the pre-RoPE KV store is designed for reuse and shared fact IDs, the architecture can be extended to support online updates (inserting, revising, or deleting facts) while maintaining consistency between the Jasper index and the KV store.
- The AI system can evolve its personalized knowledge base dynamically during a single user's session—for instance, by incrementally updating a memory fact based on a new piece of information encountered in the dialogue—without requiring a complete re-indexing or model retraining cycle.
)7. Facilitate Multi-Node and Multi-GPU Scaling:
The design separates the retrieval index (Jasper) and KV store across different GPU/CPU resources, allowing for sharding and routing requests to where specific user memories reside.
- This enables the deployment of massive, distributed personalized AI services that scale capacity and throughput beyond a single accelerator by intelligently managing memory partitioning across a cluster.
Abstract
Large language models are increasingly deployed with persistent personalized context, such as accumulated memory profiles or long conversation histories, that is shared across a user's many requests. Production memory systems (e.g., Mem0, MemGPT, and Zep) retrieve a relevant subset of this memory and inject it into the prompt, forcing the serving engine to repeatedly prefill the same content. As the retrieval budget grows, time-to-first-token (TTFT) increases even though the underlying memory is reused across requests. We present InferScale, a GPU-native LLM memory system that replaces repeated prompt prefilling with reusable KV state. InferScale precomputes each memory fact's KV representation, stores it alongside a semantic embedding on the GPU, retrieves relevant facts at serving time, and injects their KV directly into vLLM's paged cache. To support dynamically assembled memories under rotary position embeddings, we introduce Chunked RoPE, which stores keys before rotation and applies their serving-time positions during injection. However, encoding memory facts independently omits the cross-fact context available during joint prefilling. We mitigate this with Context-Window Encoding, which encodes each memory fact together with a small window of preceding conversation context while caching only the target fact's KV. InferScale is implemented through vLLM's KV-connector interface, requiring neither engine modifications nor model fine-tuning. Across three open-weight models on LoCoMo, InferScale keeps TTFT nearly constant as the retrieval budget increases: at k=50 it reduces TTFT by 72-79% (3.6-4.8x), achieves 60.3% accuracy versus 63.3% for Mem0 without serving-time recomputation, and delivers 3.7-4.5x the throughput under concurrent load. Reusable KV state thus decouples memory-conditioned serving latency from retrieved-context size while preserving application quality.
Sources
- Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory
- Gemma 2: Improving Open Language Models at a Practical Size
- LMCache: An Efficient KV Cache Layer for Enterprise-Scale LLM Inference
- GPU-Accelerated ANNS: Quantized for Speed, Built for Change
- CAGRA: Highly Parallel Graph Construction and Approximate Nearest Neighbor Search for GPUs
- MemGPT: Towards LLMs as Operating Systems
- Qwen2.5 Technical Report
- Zep: A Temporal Knowledge Graph Architecture for Agent Memory
- The Llama 3 Herd of Models
- Qwen3 Technical Report
- LazyAttention: Efficient Retrieval-Augmented Generation with Deferred Positional Encoding
Related papers
- iScheduler: Reinforcement Learning-Driven Continual Optimization for Large-Scale Resource Investment Problems
- SAMM: Sharded Automated Market Maker
- Vigil: Accountable Liveness against Selective Silence
- Steelhead: Interleaving Partially Synchronous and Asynchronous Commit Rules on a Shared DAG
- Pushing CPU Speech Synthesis to the Wall: Extreme Inference Tuning under Serverless Architecture and Billing
- SMetric: Rethink LLM Scheduling for Serving Agents with Balanced Session-centric Scheduling