CoinRAG: Contextualized Information Nugget KV Cache Reuse for Long-Context RAG

arXiv:2608.07458 · cs.CL, cs.AI, cs.IR, cs.LG · Submitted 2026-08-07 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Paper Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "CoinRAG: Contextualized Information Nugget KV Cache Reuse for Long-Context RAG".

Jane: The paper was written by Gyuwan Kim, Cheoneum Park and Tao Yang from University of California, Santa Barbara and Hanbat National University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Paper summary: ident: You're listening to the arXiv channel, where we read the latest papers so you don't have to.

Tom: Alright, everybody, welcome back. Today's paper comes from UC Santa Barbara and Hanbat National University: CoinRAG, Contextualized Information Nugget KV Cache Reuse for Long-Context RAG. If you've ever wondered why retrieval-augmented generation feels slow, this one has a clean answer.

Jane: I read it as a precision problem. RAG pulls in whole chunks of text to answer a single question, and the language model has to process all of that from scratch on every request. Most of those tokens simply aren't needed.

Lu: So the authors propose precomputing the model's key-value caches for every chunk once, offline. At query time you select tiny evidence spans — they call them nuggets — and slice their cached representations out of the precomputed caches.

Meng: And the clever part is that those nuggets aren't re-encoded in isolation. They're sliced from the cache of the full chunk, so they keep the grounding of their original document. You get a compact context without throwing away the semantics.

Tom: On the numbers it works. Under a hundred-millisecond P99 latency budget, the paper reports 41 point 7 average F1 across three multi-hop QA benchmarks, against 39 point 6 for the strongest chunk-level baseline, TurboRAG. That's a five-point-three percent relative gain, with a 1 point 84 times shorter context.

Jane: The thing that surprised me is that even with no latency limit at all, the average improvement stays above five percent. With unbounded time, you'd think the chunk-based systems would eventually win on pure recall.

Lu: Their argument is that noise is the enemy. Longer contexts don't just cost compute — they dilute attention and can actively mislead the model. A few sharply selected facts beat a wall of text.

Meng: But none of this comes free. There are real costs for offline extraction, cache storage, and model fine-tuning.

Lu: Right, and the paper is upfront about those costs in the limitations section. The offline investment is one-time per corpus and per model, not per query.

Lalam: What makes this relevant is the engineering context. Real services run under service-level agreements with tail latency targets, and a method that improves the Pareto frontier there is immediately useful to anyone operating RAG at scale.

Tom: So let's go back to page one and see how they frame that latency problem, because the whole design hangs on it.

Page 1 — The latency constraint: Jane: So we've got the big picture — RAG is slow because it re-encodes long contexts, and CoinRAG wants to shrink what the model has to see. Page one explains why that's a hard constraint rather than a nicety.

Tom: Right, and it all starts from the prefill stage, the part of inference where the model processes the prompt and retrieved documents before generating a single token. That's where the latency goes, and it happens again for every query, even when the system has served the same documents many times before.

Lu: The paper frames it through interactive services. There's a classic result in human perception that a response within about a hundred milliseconds feels instantaneous, and the authors adopt a P99 time-to-first-token budget at that level.

Jane: P99, meaning ninety-nine percent of requests have to make it under the budget, not just the average. For a service operator the tail is what users actually feel, and that's a much harder target than the mean.

Meng: Then they tie this to cache-augmented generation, the CAG paradigm. Instead of encoding documents per query, you precompute their KV caches once and reuse them. The extreme version preloads the whole knowledge base, which fixes latency but explodes memory.

Lalam: The middle ground, the one most systems settle on, is chunk-level caching. Each text chunk is cached independently, and systems use rotary position embeddings to rotate and stitch chunks back together at query time.

Tom: That RoPE rotation is the mechanical foundation of chunk-level reuse. But the paper's complaint is that a full chunk is a blunt instrument — you load tokens that have no relation to the question.

Lu: And they cite the "lost in the middle" finding, where long contexts bury the relevant evidence. That's a second argument for going finer than chunks.

Jane: So the motivation is two-sided: the latency budget forces the context to stay short, and the noise problem says short isn't just a constraint — it's actually good for accuracy.

Meng: Which naturally raises the question on page two: what should the unit of cached context be, if not the chunk?

Page 2 — Nuggets and the problem formulation: Tom: Page two answers that question with a name: information nuggets. The concept comes from information retrieval evaluation, where human assessors list the essential facts a good response ought to contain. This paper turns that into a cached, machine-usable representation.

Jane: Architecturally, context encoding moves entirely offline. Every chunk gets one forward pass, and its full KV cache is stored. At query time you never re-encode the retrieved text — you slice the cached pieces you need and assemble them into a prefix.

Lu: Formally, that prefix cache gets concatenated with a query encoding computed online. The model sees the evidence as one continuous prompt, even though it's assembled from fragments.

Meng: The coin metaphor in the title lands right here. Small pieces of value accumulated into something larger — that's literally what these nugget caches do.

Tom: The crucial design decision is that each nugget is a contiguous text span within its source chunk, recorded by start and end token indices. Those indices act as pointers into the precomputed cache, and that's what makes slicing possible.

Jane: If you'd extracted facts as free-floating sentences instead, you'd have to re-encode them from scratch and lose the document context. The span-based grounding keeps that context at zero cost.

Lu: The page also previews the three contributions — offline extraction, two-stage retrieval that selects query-relevant nuggets, and nugget-aware fine-tuning so the model learns to handle stitched-together contexts.

Lalam: There's a walkthrough in the appendix that makes it concrete, on a HotpotQA question about who wrote the novel behind the musical The Pirate Queen. Chunk retrieval brings in several documents, nugget retrieval isolates two supporting facts, and the model answers Morgan Llywelyn.

Meng: That example captures the whole point — only a handful of tokens in those chunks actually mattered, and CoinRAG puts exactly those into the context.

Tom: So now you need machinery that finds those nuggets reliably in raw text, and that's exactly what page three covers with Algorithm 1.

Page 3 — Extraction and two-stage retrieval: Jane: Page three gets into the plumbing. First comes Algorithm 1, the nugget extraction pipeline — how do you get clean, well-grounded nuggets out of raw passages?

Tom: You start with a passage and an LLM extractor that proposes candidate nuggets. The paper later reports about seven candidates per passage on average. But an extractor won't always quote text verbatim, so each candidate has to be matched back to an actual span.

Lu: That happens in stages. First an exact substring match, and if that fails, a fuzzy match against word-level spans of the passage, accepting the best match only when its similarity clears a threshold. Candidates that can't be matched to anything get dropped.

Meng: Keeping the span grounded in the source text is what makes the caching story work, because the span boundaries translate directly into KV cache positions.

Tom: Then at query time there's two-stage retrieval. A dense retriever fetches the top chunks from the corpus, and then the pre-extracted nuggets inside those chunks get ranked against the query, with the top ones selected.

Jane: That two-stage design seems natural once you see it. You're not searching a giant pool of millions of nuggets; you're searching inside documents that already looked relevant.

Lalam: There's a subtle benefit too. A nugget that scores well in isolation might actually mislead you, while one that comes from a document the retriever already trusted carries more weight.

Lu: Then comes the slicing step, where the context preservation happens. Each nugget's boundaries point into the precomputed cache of its chunk, so you slice exactly those positions. The paper stresses that the sliced states are identical to a fresh encoding of the chunk — nothing recomputed, nothing lost.

Meng: Whereas re-encoding the nugget's text on its own would compute its states without seeing the rest of the document. The ablations later show that costs several F1 points.

Tom: So — nuggets, two-stage retrieval, slicing. But the pieces come from different chunks with different positional embeddings, and you can't just glue them together. Page four takes on that alignment problem directly.

Page 4 — Position alignment and fine-tuning: Tom: Page four opens with the glue: position alignment. Every cached nugget carries rotary position embeddings from its original location, so naive concatenation would scramble the model's sense of order.

Jane: The fix is a rotation operator that shifts a cached block's positional indices by a delta, and RoPE makes that possible without re-encoding anything.

Lu: The delta for each nugget comes from a running sum — the system prompt length plus all the nugget lengths already packed in front of it. You preserve the original document order, pack the nuggets contiguously, and mask out everything in between.

Meng: Contiguous packing is what keeps total context minimal. And a shorter context isn't only about latency — it means smaller KV cache memory and faster decoding, with no change to the model architecture.

Lalam: It's genuinely unusual. The model perceives a single dense prompt, but that prompt is assembled from fragments scattered across the corpus. The position rotation makes the composition invisible to the model.

Tom: Then there's the training side. Language models are trained on continuous text, so this stitched composition creates a gap between training and inference.

Jane: Nugget-aware fine-tuning closes the gap. They build training instances by fetching relevant nuggets, composing the prefix cache exactly as at test time, and optimizing standard next-token prediction on the answer.

Lu: And the approach works even without fine-tuning — the training is calibration rather than a requirement. That lowers the barrier for adoption considerably.

Meng: I also appreciate that the training data mixes ground-truth evidence with distractors, mirroring real retrieval. You're teaching the model to work with noise, not just clean paragraphs.

Tom: So the design is complete: offline extraction, two-stage retrieval, cache slicing, position alignment, fine-tuning. Page five steps back and puts the whole package next to the existing RAG paradigms.

Page 5 — Positioning against prior work: Jane: Page five maps the design space, and the comparison is structural rather than a tuning contest. That's the useful part.

Tom: Standard RAG computes everything online — no KV reuse at all. TurboRAG represents the chunk-level caching family: each chunk is precomputed and reused, but the unit of retrieval and composition remains the whole chunk.

Lu: CacheBlend takes a middle route. It reuses caches but selectively recomputes a small subset of tokens online to restore cross-attention with preceding context — a direct attempt to fix the cross-chunk blind spot.

Meng: And KVLink inserts trainable link tokens whose caches attend to earlier chunks during encoding. So it also restores cross-chunk interaction, but through learned parameters rather than recomputation.

Tom: All four of those operate on chunks. CoinRAG's retrieval unit is the nugget, and that change cascades — shorter contexts, less noise, lower per-query latency.

Jane: The contrast with nugget-based RAG is just as sharp. GINGER and Crucible construct nuggets online per query with an LLM call, and their nuggets are free-standing text with no surrounding context.

Lu: CoinRAG flips both properties: extraction happens offline and query-independently, and each nugget remains a grounded span of its source document. Both properties are exactly what enable KV cache reuse.

Lalam: Which is why those systems don't appear in the experiments. They don't precompute caches, and their online LLM calls make them slow by construction. They're solving a different problem.

Meng: The tables on this page make the field easy to read: retrieval unit, whether KV reuse exists, what gets encoded, whether training is required. A clean separation.

Tom: And from that map, the experiments take the strongest chunk-level systems — TurboRAG, CacheBlend, KVLink — plus Standard RAG, and ask whether the nugget approach actually wins under latency budgets. That's page six.

Page 6 — Setup and main results: Jane: Page six sets up the race carefully. Three multi-hop benchmarks from LongBench — HotpotQA, 2WikiMQA, and MuSiQue — and each one requires combining evidence across documents.

Tom: Multi-hop is the right stress test because the system has to assemble nuggets from different sources into a single reasoning chain. A single-hop dataset wouldn't exercise the composition machinery at all.

Lu: The stack is concrete. GPT-4o-mini extracts nuggets offline, BGE-M3 handles retrieval, and Qwen2-7B-Instruct generates answers, chosen because its RoPE support enables the position rotation.

Meng: They also sweep retrieval counts for every method, so each point on the curves is that method's best configuration under a given budget. No cherry-picking.

Tom: Under the hundred-millisecond P99 budget, the results are consistent across all three datasets. CoinRAG scores 51 point 4 against TurboRAG's 49 point 1 on HotpotQA, 42 point 4 against 42 point 2 on 2WikiMQA, and 31 point 4 against 27 point 4 on MuSiQue.

Jane: Averaged, that's 41 point 7 versus 39 point 6 — the five-point-three percent relative gain. And the average context length is 465 tokens versus 855 for TurboRAG, so it's winning while feeding the model less than half the text.

Lalam: Better accuracy at lower cost — that's exactly what a Pareto improvement looks like. For a service operator, that combination is the most desirable result in this space.

Meng: Standard RAG is the cautionary tale in that table. It can afford exactly one retrieved chunk under the budget, and its F1 collapses. The whole caching motivation is visible in that single row.

Lu: Interestingly, KVLink and CacheBlend, which invest in cross-chunk attention, don't beat the simpler TurboRAG under this tight budget. The extra interaction costs them either latency or noise.

Tom: And that sets up page seven, where they release the latency knob entirely and test whether the advantage survives without time pressure.

Page 7 — Pareto frontiers under relaxed budgets: Jane: Page seven stress-tests the claims. They plot accuracy against latency budget, and also against context length budget — two axes of the same trade-off.

Tom: On the latency axis, CoinRAG holds the frontier up to about 116 milliseconds, where it still beats every other method. Beyond that the chunk-level systems start catching up — KVLink overtakes it on HotpotQA past roughly 160 milliseconds.

Lu: And on 2WikiMQA, given unlimited time, Standard RAG and TurboRAG edge ahead. The paper is honest about that: cross-chunk interaction can genuinely help when you have all the budget in the world.

Meng: But even with no latency limit, the three-dataset average favors CoinRAG — 42 point 7 F1 against 40 point 6 for TurboRAG. The average stays positive because removing noise wins more often than the lost interactions hurt.

Tom: The context length story is even stronger. Without a latency ceiling, CoinRAG's average length is 580 tokens against nearly four thousand for TurboRAG — a 6 point 8 times difference — and up to ten times shorter than Standard RAG in its longest settings.

Lalam: That length number is a hardware number. Live KV cache size decides how many concurrent requests fit in GPU memory, and that translates directly into serving throughput. A tenfold reduction changes the economics of a deployment.

Jane: Their interpretation convinces me: removing noise and unnecessary context offsets the loss from missing some cross-chunk interactions. Under the latency pressure of real SLAs, that trade-off tilts even further in CoinRAG's favor.

Meng: But the whole argument depends on each component of the design pulling its weight. That's exactly what the ablation studies on page eight test, one component at a time.

Page 8 — Ablations: Tom: Page eight runs the ablations, and the first one targets the slicing trick itself — the heart of the mechanism.

Jane: They compare contextualized KV slicing against re-encoding the identical nugget spans in isolation. Same retrieval, same spans, only the surrounding chunk context is missing. Peak F1 drops by 6 point 3 points on HotpotQA, 4 point 9 on 2WikiMQA, and 3 point 9 on MuSiQue.

Lu: That's direct evidence for the core claim. The cached representation carries information from the whole chunk, and that information matters when answering.

Tom: The second ablation tests two-stage retrieval against fetching nuggets directly from the whole collection in one stage. Two-stage wins by 9 point 6 to 17 point 5 percent relative in peak F1, and it selects fewer nuggets at its peak configuration.

Meng: So narrowing the candidate pool to nuggets inside already-retrieved chunks doesn't merely save compute — it improves retrieval quality. The chunk-level filter acts as a relevance prior.

Jane: Third is position alignment. Removing it, letting nuggets stay at their original positions, hurts most under a tight 75-millisecond budget, with F1 dropping 3 to 8 point 5 percent. Beyond a hundred milliseconds the penalty mostly disappears.

Tom: That pattern fits the design. Alignment buys you a compact, ordered context, and compactness is most valuable exactly when you're squeezed for budget. With room to spare, distortion matters less.

Lu: And the biggest lever is the fine-tuning. Without it, peak F1 falls by 11 point 3 points on HotpotQA, 6 point 3 on 2WikiMQA, and 6 point 4 on MuSiQue.

Meng: Which tells you the stitched context is genuinely foreign to an off-the-shelf model. It needs calibration to handle non-contiguous evidence properly.

Lalam: Every component earns its place, which is what you want from a systems paper. But the same honesty carries over to page nine, where the authors lay out the limitations of the architecture.

Conclusion: Tom: So we've reached the end, and the overall picture holds together. The paper establishes a new Pareto frontier for long-context RAG, and the gains are largest precisely under the interactive latency budgets that real services face.

Jane: And the recipe stood up to scrutiny. The offline extraction of grounded spans, the two-stage retrieval at query time, the contextual slicing of precomputed caches — plus the alignment and fine-tuning that make stitching feasible.

Lu: The trade-offs are explicit too. There are offline costs for extraction, cache storage, and fine-tuning; the cached representations are locked to a specific model checkpoint; and final quality is bounded by retrieval recall at both stages.

Meng: There's also the structural limit that nuggets from different chunks never attend to each other during encoding. That's shared with TurboRAG-style caching, and CacheBlend exists precisely to address it.

Lalam: Still, the practical case is strong. Under a hundred-millisecond P99 budget, the paper shows higher accuracy than the best chunk-level baseline while using roughly half the context tokens. For anyone serving RAG at scale, that's a meaningful win.

Tom: What impressed me most is the thoroughness — sweeping every baseline's budgets, reporting where rivals catch up, and running each component through ablations. It's a complete empirical case.

Jane: I think this one will get cited often as RAG systems push toward interactive response times.

Tom: Great conversation, everyone. Let's close the book on this paper and see what else is waiting in the arXiv queue.

Gyuwan Kim, Cheoneum Park, Tao Yang

University of California, Santa Barbara · Hanbat National University

cs.CL, cs.AI, cs.IR, cs.LG

Submitted: 2026-08-07

Updated: 2026-08-10

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 70/100

The gist: The paper introduces CoinRAG (Contextualized Information Nugget KV Cache Reuse for Long-Context RAG), a framework designed to optimize the efficiency–accuracy Pareto frontier for

Key concepts

KV cache reuse
In large language models, key-value (KV) caches store intermediate representations during generation. Reusing precomputed caches for retrieved documents avoids re-encoding them for every query, reducing latency. CoinRAG extends this by caching at the chunk level and slicing out only relevant nuggets.
Information nuggets
Small, contiguous text spans within a document that contain essential facts for answering a question. In CoinRAG, nuggets are extracted offline and stored as pointers into the chunk's KV cache, allowing precise selection at query time without losing document context.
Two-stage retrieval
A retrieval process that first uses a dense retriever to fetch top relevant chunks from a corpus, then ranks pre-extracted nuggets within those chunks against the query. This narrows the search space and leverages document-level relevance to improve nugget selection.
Position alignment
When assembling nuggets from different chunks, their rotary position embeddings must be adjusted to maintain coherent order. CoinRAG uses a rotation operator to shift positions based on a running sum, enabling contiguous packing without re-encoding.

Terminology

Summary

The paper introduces CoinRAG (Contextualized Information Nugget KV Cache Reuse for Long-Context RAG), a framework designed to optimize the efficiency–accuracy Pareto frontier for Retrieval-Augmented Generation (RAG) under low prefill latency constraints. The core motivation is stated in the abstract:

"Recent optimization studies on Retrieval-Augmented Generation (RAG) have exploited chunk-level KV cache reuse to avoid processing long retrieved contexts for higher efficiency, while significant information redundancy and noise still remain in the coarse-grained chunks."

The name CoinRAG is metaphorical:

"much like assembling small tokens (or 'coins') to accumulate a larger value, CoinRAG compositionally reuses offline-computed, fine-grained nugget caches to form a learned contextual representation efficiently in a more semantically relevant but compact manner."

Specifically, instead of full-chunk encoding, CoinRAG identifies query-relevant semantic units within retrieved chunks through two-stage retrieval and seamlessly assembles their sliced KV representations with a chunk-level context.

The main empirical claim is: "Extensive evaluations on LongBench multi-hop question answering tasks demonstrate that CoinRAG significantly reduces operational costs and outperforms the other baselines with a new Pareto frontier and an average 5.3% relative improvement in answer quality (F1) under a standard fast prefill latency budget."

The paper frames the problem around the well-known inefficiency of standard RAG at inference time. Standard RAG generates an answer to a query q by jointly encoding the system prompt p, retrieved chunks, and the query q online during the prefill stage, which introduces significant computational redundancy and high prefill latency because long contexts are repeatedly re-encoded across queries.

The authors adopt an interactive-latency framing grounded in Service Level Agreements (SLAs):

"Modern SLAs specify targets like 99% of requests under some milliseconds... a response time of 100 ms or less is perceived as instantaneous... our study adopts a Time-to-First-Token (TTFT) budget limit as the P99 latency under 100 ms. Note that lowering latency also leads to a higher throughput."

The paper notes that Standard RAG, even with KV cache reuse, is expensive in terms of TTFT. To meet SLA under a low-latency budget, it needs to limit the retrieval scope, degrading answer accuracy. CoinRAG's approach is to operat[e] RAG at a finer-grained level of information extracted in advance from text chunks with a slim representation, which also reduces the risk that relevant evidence gets lost amid long, noisy contexts.

Formally, let the corpus D be segmented into fixed-size text chunks B = bj. CoinRAG shifts encoding offline: every chunk bj ∈ B is processed through a one-time forward pass to compute and store its full-context key-value (KV) representations Cbj. At online inference, the framework dynamically extracts query-relevant textual units within retrieved chunks, slices their corresponding precomputed KV caches, and aligns them into a single prefix context cache Cctx. The final cache is constructed as:

KVCoinRAG = Cctx ⊕ KVM(q; Cctx), where ⊕ denotes KV cache concatenation along the sequence dimension, and KVM(q; Cctx) represents the query encoding pass.

CoinRAG extracts candidate nuggets from each text chunk offline, identifying each as a contiguous span defined by start and end token indices within its source chunk. This span-level grounding is crucial: This span-level grounding lets us later slice the corresponding portion of the precomputed chunk cache to form a contextualized nugget KV cache, rather than needing to re-encode the nugget's text in isolation.

Algorithm 1 proceeds as follows: an LLM first proposes candidate nuggets from passage p; each candidate is matched back to a token span. Stage A (exact match) checks whether it appears verbatim as a substring of p. If not, Stage B (fuzzy match) searches word-level spans of p near the candidate's length and match against the one with the highest similarity to the candidate, accepting it only if this similarity exceeds a threshold τ. Otherwise, Stage C (abandon) drops the candidate. The accepted spans, marked by start/end token positions, form the extracted nuggets.

At runtime, CoinRAG identifies the most informative segments via two-stage retrieval scalable to large corpora:

"Given a query q, a dense retriever first fetches the top-kc text chunks from the corpus to construct a retrieved subset Bret ⊂ B. We then evaluate and rank the pre-extracted candidate nuggets nested exclusively within these retrieved chunks Bret based on their embedding similarity to q, keeping the top-k entries."

Crucially, the boundary indices [si, ei] are used "as pointers to slice the corresponding portions of the precomputed offline chunk cache Cbi. This operation yields contextualized nugget KV caches, denoted as Cbi[si: ei], ensuring that the representations inherently retain the deep, document-level context established during full-chunk preprocessing. The paper emphasizes that every token's KV state within the sliced range [si, ei] is identical to what it would be under a fresh encoding of the chunk, carrying the same contextual grounding that full-chunk encoding provides, whereas isolated nugget encoding forfeits exactly this: re-encoding a nugget's raw text on its own recomputes its KV states from scratch, without access to the rest of the chunk."

Because individual nugget caches come from different source chunks and positions, their RoPE positional encodings mismatch. CoinRAG adapts a position rotation operator Rot(C; ∆) that shifts positional indices by an offset. The unified context cache is:

Cctx = Cp ⊕ Rot(Cb1[s1: e1]; ∆1) ⊕... ⊕ Rot(Cbk[sk: ek]; ∆k).

The strategy is order-preserving, contiguous position alignment that preserves the original document sequence order while packing nuggets consecutively to eliminate structural gaps. The rotation offset is "dynamically calculated as ∆i = p + Σ j<i(ej − sj + 1). By masking non-retrieved intermediate text and assigning continuous positional indices, this mechanism minimizes the total context length, thereby reducing GPU memory footprints and accelerating decoding without altering the model architecture."

To bridge the structural training-inference gap, as LLMs are trained on continuous sequences, CoinRAG incorporates nugget-aware fine-tuning. For each training instance (query q, target answer y, document collection with ground-truth and distractors), the pipeline fetch[es] the top-k relevant nuggets for q from these documents and dynamically construct[s] the prefix context cache Cctx via the position alignment. The objective is standard cross-entropy loss on target answer tokens.

The paper positions CoinRAG against several lines of work, formalized as:

  • Standard RAG: KVStandardRAG = KVM(p ⊕ tchunk ⊕ q) — computes everything online, suffering severe prefill latency due to full online forward propagation over long sequences for every query.

  • Standard CAG / TurboRAG: KVStandardCAG = Cchunk ctx ⊕ KVM(q; Cchunk ctx) — caches full retrieved chunks offline; TurboRAG and Block-attention precomput[e] and reus[e] chunk-level KV caches to isolate cross-chunk attention and slash TTFT.

  • CacheBlend: reuses precomputed KV caches while selectively recomputing a small subset of tokens to restore cross-attention with preceding context.

  • KVLink: inserts trainable link tokens into each chunk, whose KV states are computed by attending over the previous chunks to restore cross-chunk self-attention.

  • CoinRAG: encodes each chunk independently offline, but at the finer granularity of nuggets rather than full chunks, which reduces latency.

Regarding nugget-based RAG, the paper distinguishes CoinRAG from GINGER and Crucible: "Nugget extraction in CoinRAG occurs offline and is query-independent, becoming query-specific only after online two-stage retrieval, and each nugget is a text span from the original document that recovers the broader context of its source chunk during inference. GINGER and Crucible instead construct nuggets online per query with an LLM, and each stands alone with no surrounding context. Because those methods do not study KV cache precomputation and instead incur high prefill latency overhead from repeated expensive LLM calls during query processing," they are excluded from the empirical comparison.

Figure 1 illustrates the key architectural difference: CoinRAG's prefill context is assembled from offline-computed nugget slices, whereas Standard RAG computes all KV caches online, and CacheBlend/TurboRAG/KVLink operate at the chunk level.

  • Baselines: Standard RAG, TurboRAG, CacheBlend, KVLink. (Block-attention is omitted since TurboRAG represents it.)

  • Metrics: Token-level F1-score for accuracy; TTFT latency (ms) for efficiency, the time to construct the prefix context and process the query during the prefill stage; plus prefill length in tokens as a direct proxy for KV cache memory usage.

  • Datasets: Three multi-hop multi-document QA benchmarks from LongBench: HotpotQA, 2WikiMQA, MuSiQue, that demand cross-document aggregation and reasoning.

  • Models: GPT-4o-mini for offline nugget extraction; BGE-M3 as the dense retriever; Qwen2-7B-Instruct as the base LLM, chosen for its native support for flexible position rotation via RoPE.

  • Sweeps: chunk-based methods sweep kc ∈ 1, 2, 3, 4, 5, 10, 20, 30, 50; CoinRAG sweeps kc then k ∈ 1, 2, 3, 4, 5, 10, 20, 30, 50, 100, 150, 200; CacheBlend additionally sweeps recomputation ratio r ∈ 0.0, 0.1,..., 1.0.

  • Fine-tuning: One epoch over 277,280 training instances from IRCoT training splits, with hyperparameters detailed in Appendix G (LR 5×10−6, cosine scheduler, warmup ratio 0.03, micro-batch 1, gradient accumulation 16, max sequence length 1024, training chunks kc=5, training nuggets k=10).

"Standard RAG is forced only to retrieve 1 document to meet the budget, while other baselines offer more flexibility with KV cache reuse. In all 3 datasets, CoinRAG outperforms Standard RAG, CacheBlend, TurboRAG, and KVLink, while TurboRAG is the strongest among the other baselines."

  • The F1 score of CoinRAG is 5.3% higher than TurboRAG averaged over 3 datasets, while the contextual token length of CoinRAG is 1.84× shorter.

  • Per-dataset F1 improvements over TurboRAG: 4.7% (HotpotQA), 0.5% (2WikiMQA), 14.6% (MuSiQue).

"When removing the latency limit entirely, TurboRAG remains the strongest other baseline, and CoinRAG outperforms it in 2 out of 3 datasets with a 3-dataset average improvement of 5.2% (42.7 vs. 40.6) and with a 6.8× shorter average length."

The paper explains the trade-off: "Although the slim KV cache representation in CoinRAG can sometimes miss useful interactions across chunks, on average this design can still offset the loss with a larger gain by removing noise and unnecessary context that misleads or slows down LLM inference."

Figure 2 plots F1 versus P99 TTFT latency budget; Figure 3 plots F1 versus prefill context length limit. "CoinRAG establishes a stronger empirical Pareto frontier, maximizing answer quality while substantially reducing hardware resource usage. This smaller KV cache footprint per query also allows more concurrent requests to fit in GPU memory, translating into higher serving throughput for RAG systems under low-latency SLAs. Notably, Even with no length limit, CoinRAG's context length remains up to 10.1× shorter than Standard RAG's, the longest among all baselines under this setting."

Four ablations are evaluated on all three datasets (Figure 4):

  1. Contextualized vs. Isolated Nugget Encoding: "encoding nuggets in isolation lowers peak F1-score by 6.3, 4.9, and 3.9 points on the 3 datasets. This confirms that the cached KV representations carry contextual information beyond what the isolated nugget span alone conveys."

  2. Two-Stage vs. One-Stage Retrieval: "Two-stage retrieval reaches a higher peak F1-score on all 3 datasets, a relative gain of 9.6 to 17.5 percent over one-stage retrieval... This confirms that narrowing the candidate pool to nuggets within already-retrieved chunks improves retrieval quality."

  3. Position Alignment: "Position alignment is most beneficial under a tight 75 ms P99 budget, improving F1 by 3.0–8.5% across the 3 datasets. From a 100 ms P99 budget onward, position alignment achieves a comparable F1 score to no position alignment."

  4. Nugget-Aware Fine-Tuning: "Nugget-aware fine-tuning improves peak F1-score by +11.3 points on HotpotQA, +6.3 points on 2WikiMQA, and +6.4 points on MuSiQue. This confirms that nugget-aware fine-tuning bridges the training-inference alignment gap by calibrating the model against position embedding shifts induced by stitching together non-contiguous nugget spans."

The paper concludes: "Under a standard 100 ms P99 latency budget, CoinRAG outperforms the F1 score of the best competitor, TurboRAG, by 4.7%, 0.5%, and 14.6% for 3 datasets, with an average gain of 5.3%. Without a latency limit, CoinRAG still outperforms the baselines on average for 3 datasets by 5.2% F-1 improvement or more, indicating that the gain by removing noise and unnecessary context by CoinRAG outweighs the loss due to missing some interaction. Position alignment is most beneficial when the P99 latency budget is tight (e.g., under 75 ms), and remains comparable to no position alignment beyond that budget."

The paper explicitly states several constraints:

  • Offline costs: "pre-encoding text into static KV representations significantly accelerates online prefill runtimes, it introduces offline resource costs that scale with the corpus size, including disk storage overhead... and a non-trivial one-time nugget-aware fine-tuning cost. These are acceptable assuming documents do not change frequently." The appendix quantifies: 265–325 for training-corpus nugget extraction, under 2 for evaluation corpora; offline cache construction under 30 minutes total (27–28 MB per chunk: 146 GB HotpotQA, 75 GB 2WikiMQA, 173 GB MuSiQue); 140 GPU-hours for fine-tuning on one RTX PRO 6000.

  • Model coupling: The precomputed cache layer is inherently coupled to the specific model checkpoint... updating the backbone language model or altering its architecture forces a re-encoding of documents.

  • Retrieval recall bound: The downstream answer quality remains bounded by the initial chunk-level and nugget-level retrieval recall.

  • No cross-chunk attention: nuggets originating from different chunks do not attend to each other during encoding, a limitation shared with TurboRAG-style chunk-level caching.

  • No inter-query cache reuse: the evaluation does not consider inter-query KV-cache reuse when some consecutive queries may share the same retrieved documents, a setting shared with prior work.

The appendices additionally include a step-by-step inference walkthrough (Appendix D), prompt templates (Appendix B), extraction statistics showing mean nugget length is far shorter than mean passage length with low overlap ratios (Appendix C), detailed hardware/I/O infrastructure (Appendix F), fine-tuning hyperparameters (Appendix G), and categorized success/failure inference examples (Appendix H), including cases where the model is misled by surface-level entity frequency or where retrieval fails to surface the canonical evidence span.

Improvements for AI systems

Specific improvements to AI systems based on CoinRAG:

  1. Replace full-chunk KV cache reuse with fine-grained, query-relevant nugget KV cache reuse.

Instead of caching and attending to entire retrieved chunks, the system precomputes KV representations for semantic spans (nuggets) offline, then slices and assembles only the query-relevant spans at inference time.

Improved capability: The system can serve RAG under tight latency budgets (e.g., P99 < 100 ms) while maintaining higher answer quality—outperforming full-chunk caching baselines by up to 5.3% average F1 and using 1.84× shorter context lengths under that budget.

  1. Two-stage retrieval (chunk-level then nugget-level) for evidence selection.

First retrieve top-kc chunks with a dense retriever, then rank pre-extracted nuggets within those chunks by embedding similarity to the query, keeping top-k.

Improved capability: The system improves retrieval precision for multi-hop QA, achieving 9.6–17.5% relative peak-F1 gains over one-stage nugget retrieval, while limiting the candidate pool to avoid noise from irrelevant corpus regions.

  1. Contextualized nugget encoding via slicing precomputed chunk caches.

The system reuses the KV states from the original full-chunk encoding for a nugget's span, instead of re-encoding the nugget text in isolation.

Improved capability: The system retains document-level context within each nugget, preventing accuracy degradation (6.3, 4.9, and 3.9 F1 points on HotpotQA, 2WikiMQA, and MuSiQue, respectively) and avoiding recomputation latency.

  1. Position-aware KV cache composition with RoPE rotation offsets.

When stitching nugget slices from different source chunks, the system applies positional rotation offsets to align positions contiguously, eliminating structural gaps and reducing total context length.

Improved capability: The system packs more relevant evidence into a shorter prefix, lowering GPU memory usage and speeding up decoding, while maintaining accuracy especially under strict latency limits (3.0–8.5% F1 improvement at a 75 ms P99 budget).

  1. Nugget-aware fine-tuning to close the training–inference gap.

The system is fine-tuned on training examples constructed exactly as at inference: top-k nuggets retrieved from a document collection, assembled into a combined KV cache with position alignment, then trained with cross-entropy loss on the target answer.

Improved capability: The model adapts to non-contiguous, stitched contexts, yielding large F1 gains (+11.3, +6.3, and +6.4 points on the three datasets) and making the inference-time cache composition robust.

  1. Offline precomputation of chunk/nugget KV caches to shift prefill cost out of the critical path.

Documents are encoded once offline; online prefill only encodes the query and concatenates the precomputed caches.

Improved capability: The system achieves near-instantaneous prefill, dramatically reducing Time-to-First-Token and enabling more concurrent requests to fit in GPU memory (as the per-query KV footprint is dramatically smaller—up to 10.1× shorter than standard RAG without a latency limit).

  1. Use of single-pass, query-independent nugget extraction with span grounding.

Nuggets are extracted offline by an LLM, exact-matched or fuzzy-matched back to token spans, and indexed for later slicing.

Improved capability: The system supports scalable, reusable evidence indexing: the same offline caches serve arbitrary future queries, and the grounded spans serve as pointers into the original chunk cache without requiring online LLM calls.

What the improved AI system can do:

  • Answer multi-hop questions that require combining evidence across multiple documents, with higher accuracy than standard RAG and chunk-cache RAG systems, under the same latency constraints.

  • Operate within strict Service Level Agreements (e.g., 99% of requests under 100 ms time-to-first-token) while using far less GPU memory and fewer prefill tokens.

  • Serve high-throughput RAG applications by packing more requests per GPU, since each request's prefix context is composed of only the most relevant, compact nugget slices rather than bulky full chunks.

  • Maintain answer quality even when the total context window is severely limited, because the assembled cache contains less noise and more semantically relevant evidence.

  • Adapt to non-contiguous context layouts through fine-tuning, making the system robust to compressed, reordered, or selectively packed contexts.

Sources

Related papers