Working Around the Compute Ceiling: Byte-Exact Memory in Galahad Makes LLM Reading a One-Time Cost LLM Reading a One-Time Cost
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Working Around the Compute Ceiling".
Jane: A transformer language model performs a bounded amount of computation per token, and recent work by Vishal Sikka, former CEO of Infosys,
Tom: First, who's behind it and why it matters.
Title and authors: Tom: Moving on to the actual summary of "Working Around the Compute Ceiling: Byte-Exact Memory in Galahad Makes LLM Reading a One-Time Cost," they explain that instead of recomputing everything from scratch every time, Galahad uses two main parts: Taliesin for saving the model’s key–value state for specific text blocks, and Blaise to organize the actual documents into byte-exact sections on the CPU.
Jane: So, to put it simply, they are splitting things up: one part manages what the AI has actually "thought" about—the KV state—and another part manages where all that raw text lives in a structured way so it can be quickly retrieved. This separation is what allows for that idea of a one-time reading cost to happen.
Lu: The diagram shown in Figure one demonstrates how Blaise selects the section of text a question needs, and then Taliesin supplies the saved state if the model has already processed those exact bytes, which is a very clever way to manage memory across different stages of processing.
Meng: It sounds like they are trying to decouple the initial reading process from the actual computation process so we don't have to pay for reading a huge corpus every single time we ask a follow-up question. That’s a pretty big structural change for how we deploy these systems.
Lalam: For our culture, this means an assistant can handle hundreds of repeated queries against our internal manuals without slowing down or using up resources unnecessarily, making it feel much more responsive and intelligent when users keep asking the same things.
The paper's summary: Tom: Now let’s look at what they actually improved in this work. The main improvement is that they moved away from a system where every second question recomputed everything to one that uses Taliesin for KV state storage and Blaise for document section organization, creating a mechanism where reading text becomes a one-time cost.
Jane: They also built in some really strict safety checks, which is super important because they stressed that approximate reuse can fail silently if you aren't careful. They’ve included checks like saving and restoring two hundred sixty-two thousand one hundred forty-four of two hundred sixty-two thousand one hundred forty-four logits bit-identical to make sure the memory never leads to a wrong answer.
Lu: The design rule they follow is really interesting: "serve without memory, never with wrong memory," which means if any record fails a check on load, the system doesn't just guess; it goes back and recomputes the block and completes normally. That’s a strong commitment to correctness over speed when those critical checks are involved.
Meng: From an engineering viewpoint, that safety mechanism is what gives me confidence because we need assurance that when we offload state to the CPU or disk, we have these cryptographic guarantees ensuring the integrity of what's being loaded back into memory.
Lalam: That level of verification gives us peace of mind because it means the system isn't just guessing based on a lucky guess; it’s actively confirming that the information loaded is exactly what was stored before it goes back to the model.
The paper's improvements: Tom: So, wrapping up this discussion on "Working Around the Compute Ceiling: Byte-Exact Memory in Galahad Makes LLM Reading a One-Time Cost LLM Reading a One-Time Cost," the main implication is that we can shift inference costs from being tied to total text processed to being tied only to new information read. This means repeated work becomes nearly free, provided the data matches what’s already stored in Taliesin and Blaise.
Jane: That’s right, Tom. The paper suggests that for any workload that returns to the same documents, we should structure our systems so that the cost scales with how much *new* text is introduced, not by how much total text was ever processed.
Lu: I think the biggest potential here is enabling smaller models to effectively handle huge corpora because they don't need massive context windows if they can leverage this persistent memory layer to access only the most relevant parts.
Meng: We should keep an eye on how this stateful inference moves us toward building more cost-aware pipelines where the system dynamically chooses between loading saved memory or performing a full recomputation based on input similarity.
Lalam: For our team, this means we can build tools that are incredibly fast for routine tasks because the most common queries will be served instantly from the memory layer, freeing up resources for more complex reasoning.
Tom: That’s it for this session on Galahad; it really shows how precise memory management can fundamentally change how we think about scaling AI services. We’ll catch you next time as we look at something completely different in the research world.
Conclusion: Tom: So we’ve been diving deep into Galahad today, looking at how they managed to make reading text a one-time cost by using Taliesin and Blaise for byte-exact memory.
Jane: That’s right, Tom; essentially, they solved the problem of paying for repeated computation by treating stored data as a persistent asset rather than a fresh calculation every time.
Lu: I think the real power here is how that byte-exact matching works across different runtimes, which suggests we can achieve very stable state reuse even when moving between systems.
Meng: I see the practical value in that stability; if we can guarantee correctness through those checks, we get a much more predictable inference cost for our production pipelines.
Lalam: For us, this means our AI tools can become so much more responsive and reliable for our users because they won't be bogged down by repeated loading times when they ask the same questions.
Tom: Exactly; the paper really shows how precise memory management can fundamentally change how we think about scaling AI services.
Jane: It’s a fascinating look at moving inference from a stateless computation model to one with persistent memory, which is exactly what we need for many complex tasks.
Lu: We should definitely keep an eye on how this stateful approach interacts with other multimodal research, maybe seeing how it handles the complexity of visual or textual inputs in future work.
Meng: I’m curious if this memory persistence can be leveraged to reduce the sheer size of the models we have to deploy on edge devices.
Lalam: I hope we see this kind of efficiency bleeding into our core applications, making them feel much more intelligent and helpful every single time a user interacts with them.
Tom: Well, that wraps up our deep dive into "Working Around the Compute Ceiling: Byte-Exact Memory in Galahad Makes LLM Reading a One-Time Cost." We’ve seen how memory can be structured to reduce the cost of reading text significantly.
Jane: It was certainly an interesting look at how to manage computational budgets by focusing on reusing existing work.
Lu: Next week, we’re looking at something totally different, so get ready for a shift in focus as we explore new areas in multimodal reasoning and image-text alignment.
Sietse Schelpe
Corbenic AI
cs.CL, cs.AI, cs.IR, cs.LG, cs.PF
Submitted: 2026-09-30
Updated: 2026-09-30
Code: https://github.com/corbenicai/galahad
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 92/100
The gist: A transformer language model performs a bounded amount of computation per token, and recent work by Vishal Sikka, former CEO of Infosys, argues that this bound limits which tasks a model can carry
Key concepts
- Taliesin
- This memory component stores the model's key-value (KV) state for a specific block of text. It enables reuse by loading this saved state when the exact same text appears again, removing repeated computation for previously processed data.
- Blaise
- Blaise organizes documents into byte-exact sections on the CPU. Instead of feeding the entire corpus, Blaise selects only one relevant section for a question, significantly reducing the amount of text read by the model during inference.
- Compute Ceiling
- This refers to the limit imposed by a transformer model's computational budget per token. Galahad works around this ceiling by keeping repeated work and verification away from the model, allowing it to handle large corpora with exact memory.
- Exactness
- The system prioritizes exactness because approximate reuse fails silently. Galahad treats any mismatch in stored data as a failed load, forcing the model to recompute. This ensures that memory never changes an answer.
Terminology
Summary
A transformer language model performs a bounded amount of computation per token, and recent work by Vishal Sikka, former CEO of Infosys, argues that this bound limits which tasks a model can carry out or verify. This paper addresses how much of that computational budget is spent on useful work versus repeating work the model has already done.
The gist
Galahad is a memory layer for vLLM, SGLang and llama.cpp that makes reading text a one-time cost by storing the model’s key–value (KV) state and document sections, allowing subsequent requests containing the same data to load this state instead of recomputing it.
System Architecture
Galahad is implemented as a shared library that plugs into three production runtimes: vLLM as a KV connector, SGLang as a HiCache storage backend, and llama.cpp through slot save and restore. The system utilizes two primary memories: Taliesin stores the model’s own KV state for a block of text, while Blaise keeps the documents themselves organized into sections on the CPU.
(Figure 1 illustrates this flow: Blaise selects the section of text a question needs; Taliesin supplies saved KV state for any block the model has already processed, and saves new blocks after their first prefill.)
Memory Mechanisms
The system achieves reuse through two distinct methods:
-
Taliesin removes repeated computation by storing the model’s own KV state for a block of text and loading it when the same bytes appear again. The design rule is
serve without memory, never with wrong memory: if a record is missing, damaged, or fails a check, the runtime recomputes the block and the request completes normally.
Reuse happens atblock level,
wheretoken positions match exactly.
-
Blaise removes unnecessary reading by storing each document as byte-exact text divided into sections. For a question, Blaise selects one section on the CPU and passes that section’s text to the model, which then reads it and answers. This allows for
reading less,
with Blaise passing about 668 tokens per question instead of the full corpus in a recall test.
Correctness and Verification
Exactness is paramount because approximate reuse fails silently.
Galahad treats any mismatch as a failed load and recomputes, ensuring that memory never changes an answer.
The system incorporates several safety mechanisms to ensure correctness:
(Table 1 summarizes the correctness checks, including saving and restoring 262,144 of 262,144 logits bit-identical.)
The paper details sabotage testing where each safety mechanism (confirmation hash, licence signature, tenant separation, byte comparison on load) was tested three times: with the defense on, with it switched off, and with it restored. A defence is effective only if the attack succeeds when it is off and fails when it is on.
Furthermore, analysis tools found one out-of-bounds read which was subsequently fixed before release.
Evaluation Results
The evaluation across seven real-world datasets showed significant gains:
(Table 2 compares accuracy against median time per question for the recall test.)
Taliesin alone (the model reads the whole corpus, loaded instead of recomputed) answered 98 to 100 questions on the three runtimes. On llama.cpp, this resulted in a 3.1× faster
response time than the baseline without Galahad and 79% less GPU energy.
With Blaise added, the model achieved 100 of 100 answers on every runtime with a median time of 0.64 seconds and 213 Joules per question on vLLM.
Implications for Inference Cost
The results support that reuse turns most prompt computation into loading.
For workloads returning to the same documents, the cost grows with the text the model has not read before.
This shift moves inference from a stateless computation to one with persistent memory, where a model should pay to read a document once, and every later question about it should cost only the question and the answer.
Galahad works around the compute ceiling by keeping repeated work, search, and verification away from the model. The main implication is that this stateful inference can enable smaller models with exact memory of a large corpus to answer questions that would otherwise require a model with a much longer context window.
Availability
Galahad is available as a free, non-commercial beta for one GPU at https://github.com/corbenicai/galahad, and commercial pilots are available on request. The system ships as one shared library (libgalahad.so) for Linux x86-64 with two runtime dependencies (OpenSSL’s libcrypto and zstd).
Improvements for AI systems
Here are the specific improvements that can be made to existing AI systems based on this research, along with a description of what those improved systems can achieve:
The core improvement is shifting LLM inference from a stateless, recomputational process to a stateful, memory-persistent process by introducing the Galahad framework. This fundamentally changes the unit of inference cost from total text processed
to new text read.
Here are specific improvements categorized by their technical mechanism:
-
Implement a shared memory layer that plugs into existing serving runtimes (vLLM, SGLang, llama.cpp) as a single library (libgalahad.so). This allows for stateful persistence across requests without modifying model weights or runtime source code directly.
-
Incorporate the Taliesin memory component to store the model’s Key-Value (KV) state for specific blocks of prompt text on disk, indexed by byte-exact input fingerprints (model ID + tenant ID + input bytes). This allows subsequent requests containing identical byte sequences to load the KV state instead of recomputing it.
-
Implement the Blaise memory component to store the corpus as byte-exact text, pre-divided into sections. When a question is posed, Blaise selects only the necessary section (e.g., 668 tokens) from the corpus and passes only that exact text to the model, rather than feeding it a massive context window or recomputing the entire prompt.
-
Enforce strict correctness checks on all loaded state using cryptographic hashes and byte comparisons before loading, ensuring that restored state is bit-identical to the freshly computed state (e.g., 262,144 logits match). Any load failure triggers a recomputation, guaranteeing correctness over approximation.
The improved AI systems can achieve the following specific capabilities:
-
Support "Persistent Q&A Agents" for high-frequency tasks (e.g., customer support assistants or coding agents) where the same documents, code snippets, or agreements are queried repeatedly. These agents will drastically reduce latency and energy consumption because repeated queries about known information cost only the question/answer time rather than the full prompt prefill cost.
-
Enable
Knowledge-Intensive RAG Systems
that operate on large corpora (millions of tokens) without incurring massive per-question costs or requiring excessively long context windows. The system can selectively retrieve and pass only the 668 most relevant tokens, leading to significant reductions in time (up to 14.5x faster for Blaise mode) and GPU energy compared to feeding the entire corpus or a truncated window. -
Achieve
Model Agnostic State Reuse
across different hardware and model architectures (e.g., loading KV state saved on an A6000 onto an RTX 4090) due to the strict byte-identity checks, allowing for cross-platform state portability. -
Create
Cost-Aware Inference Pipelines
that dynamically determine whether to recompute or load memory based on input similarity, moving inference from a purely stateless computation to one where the cost scales with the novelty of the information requested, not the total text processed in the prompt.
Abstract
A transformer language model performs a bounded amount of computation per token, and recent work by Vishal Sikka, former CEO of Infosys, argues that this bound limits which tasks a model can carry out or verify (arXiv:2507.07505). We ask how much of the budget beneath that ceiling is spent on work the model has already done. Serving is stateless across requests: a model that answers a second question about a document recomputes the document's attention state from the first token. On seven real-world datasets, 98.7% of prompt tokens were text the model had already read. We present Galahad, a memory layer for vLLM, SGLang and llama.cpp that makes this reading a one-time cost. Taliesin saves the model's key-value (KV) state for a block of text and loads it on the next request that contains the same bytes, instead of recomputing it. Blaise keeps the documents themselves and passes the model only the section a question needs. On a recall test with 100 facts hidden in a 97,000-token corpus (Gemma 4 31B), Taliesin alone let the model attend to the whole corpus and answered 98 of 100 on llama.cpp at 3.0 s and 572 J per question, against 10 of 100, 9.3 s and 2,754 J for the same model without Galahad, which could hold only the last 12,000 tokens. With Blaise added, the model read about 668 tokens per question and answered 100 of 100 on all three runtimes at 0.59-0.64 s and 200-213 J; a tuned RAGFlow pipeline answered 77. Storing the corpus is a one-time cost of about 100 s and 28 kJ, whose energy is recovered after 13 questions. Restored state is bit-identical: all 262,144 output logits matched after restart, rehydration and hot-load. Galahad worked with all 30 models we tested under vLLM, and it fails closed: any load that does not pass its checks is recomputed. Together these results move LLM serving from stateless to stateful inference.
Sources
- Cartridges: Lightweight and general-purpose long context representations via self-study
- Prompt Cache: Modular Attention Reuse for Low-Latency Inference
- RAGCache: Efficient Knowledge Caching for Retrieval-Augmented Generation
- Efficient Memory Management for Large Language Model Serving with PagedAttention
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
- Lost in the Middle: How Language Models Use Long Contexts
- CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving
- LMCache: An Efficient KV Cache Layer for Enterprise-Scale LLM Inference
- Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
- A Frozen 12B Beats Frontier Models on Verified Work: 100% Accuracy, 0 Tokens, Bit-Exact, Forever
- Smarter and Cheaper at Once: Byte-Exact KV-Cache Grafting Turns a Frozen Small Model into a Verified-Knowledge Flywheel
- Merlin: Deterministic Byte-Exact Deduplication for Lossless Context Optimization in Large Language Model Inference
- Byte-Exact Deduplication in Retrieval-Augmented Generation: A Three-Regime Empirical Analysis Across Public Benchmarks
- Hallucination Stations: On Some Basic Limitations of Transformer-Based Language Models
- CacheBlend: Fast Large Language Model Serving for RAG with Cached Knowledge Fusion
- SGLang: Efficient Execution of Structured Language Model Programs
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering