Working Around the Compute Ceiling: Byte-Exact Memory in Galahad Makes LLM Reading a One-Time Cost LLM Reading a One-Time Cost

summary

Video file (mp4)

The gist

A transformer language model performs a bounded amount of computation per token, and recent work by Vishal Sikka, former CEO of Infosys, argues that this bound limits which tasks a model can carry

In short

Galahad is a memory layer for LLMs that stores model data and key-value states to make reading text a one-time cost. It uses two memories, Taliesin for KV state reuse and Blaise for document section storage. This allows models to answer repeated questions about the same data without recomputing everything, shifting inference from computation to loading.

Key concepts

Taliesin
This memory component stores the model's key-value (KV) state for a specific block of text. It enables reuse by loading this saved state when the exact same text appears again, removing repeated computation for previously processed data.
Blaise
Blaise organizes documents into byte-exact sections on the CPU. Instead of feeding the entire corpus, Blaise selects only one relevant section for a question, significantly reducing the amount of text read by the model during inference.
Compute Ceiling
This refers to the limit imposed by a transformer model's computational budget per token. Galahad works around this ceiling by keeping repeated work and verification away from the model, allowing it to handle large corpora with exact memory.
Exactness
The system prioritizes exactness because approximate reuse fails silently. Galahad treats any mismatch in stored data as a failed load, forcing the model to recompute. This ensures that memory never changes an answer.

Terminology used across episodes

This episode discusses

The paper

Working Around the Compute Ceiling: Byte-Exact Memory in Galahad Makes LLM Reading a One-Time Cost LLM Reading a One-Time Cost · Read on arXiv

Sietse Schelpe

Corbenic AI

A transformer language model performs a bounded amount of computation per token, and recent work by Vishal Sikka, former CEO of Infosys, argues that this bound limits which tasks a model can carry out or verify (arXiv:2507.07505). We ask how much of the budget beneath that ceiling is spent on work the model has already done. Serving is stateless across requests: a model that answers a second question about a document recomputes the document's attention state from the first token. On seven real-world datasets, 98.7% of prompt tokens were text the model had already read. We present Galahad, a memory layer for vLLM, SGLang and llama.cpp that makes this reading a one-time cost. Taliesin saves the model's key-value (KV) state for a block of text and loads it on the next request that contains the same bytes, instead of recomputing it. Blaise keeps the documents themselves and passes the model only the section a question needs. On a recall test with 100 facts hidden in a 97,000-token corpus (Gemma 4 31B), Taliesin alone let the model attend to the whole corpus and answered 98 of 100 on llama.cpp at 3.0 s and 572 J per question, against 10 of 100, 9.3 s and 2,754 J for the same model without Galahad, which could hold only the last 12,000 tokens. With Blaise added, the model read about 668 tokens per question and answered 100 of 100 on all three runtimes at 0.59-0.64 s and 200-213 J; a tuned RAGFlow pipeline answered 77. Storing the corpus is a one-time cost of about 100 s and 28 kJ, whose energy is recovered after 13 questions. Restored state is bit-identical: all 262,144 output logits matched after restart, rehydration and hot-load. Galahad worked with all 30 models we tested under vLLM, and it fails closed: any load that does not pass its checks is recomputed. Together these results move LLM serving from stateless to stateful inference.

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Working Around the Compute Ceiling".

Jane: A transformer language model performs a bounded amount of computation per token, and recent work by Vishal Sikka, former CEO of Infosys,

Tom: First, who's behind it and why it matters.

Title and authors: Tom: Moving on to the actual summary of "Working Around the Compute Ceiling: Byte-Exact Memory in Galahad Makes LLM Reading a One-Time Cost," they explain that instead of recomputing everything from scratch every time, Galahad uses two main parts: Taliesin for saving the model’s key–value state for specific text blocks, and Blaise to organize the actual documents into byte-exact sections on the CPU.

Jane: So, to put it simply, they are splitting things up: one part manages what the AI has actually "thought" about—the KV state—and another part manages where all that raw text lives in a structured way so it can be quickly retrieved. This separation is what allows for that idea of a one-time reading cost to happen.

Lu: The diagram shown in Figure one demonstrates how Blaise selects the section of text a question needs, and then Taliesin supplies the saved state if the model has already processed those exact bytes, which is a very clever way to manage memory across different stages of processing.

Meng: It sounds like they are trying to decouple the initial reading process from the actual computation process so we don't have to pay for reading a huge corpus every single time we ask a follow-up question. That’s a pretty big structural change for how we deploy these systems.

Lalam: For our culture, this means an assistant can handle hundreds of repeated queries against our internal manuals without slowing down or using up resources unnecessarily, making it feel much more responsive and intelligent when users keep asking the same things.

The paper's summary: Tom: Now let’s look at what they actually improved in this work. The main improvement is that they moved away from a system where every second question recomputed everything to one that uses Taliesin for KV state storage and Blaise for document section organization, creating a mechanism where reading text becomes a one-time cost.

Jane: They also built in some really strict safety checks, which is super important because they stressed that approximate reuse can fail silently if you aren't careful. They’ve included checks like saving and restoring two hundred sixty-two thousand one hundred forty-four of two hundred sixty-two thousand one hundred forty-four logits bit-identical to make sure the memory never leads to a wrong answer.

Lu: The design rule they follow is really interesting: "serve without memory, never with wrong memory," which means if any record fails a check on load, the system doesn't just guess; it goes back and recomputes the block and completes normally. That’s a strong commitment to correctness over speed when those critical checks are involved.

Meng: From an engineering viewpoint, that safety mechanism is what gives me confidence because we need assurance that when we offload state to the CPU or disk, we have these cryptographic guarantees ensuring the integrity of what's being loaded back into memory.

Lalam: That level of verification gives us peace of mind because it means the system isn't just guessing based on a lucky guess; it’s actively confirming that the information loaded is exactly what was stored before it goes back to the model.

The paper's improvements: Tom: So, wrapping up this discussion on "Working Around the Compute Ceiling: Byte-Exact Memory in Galahad Makes LLM Reading a One-Time Cost LLM Reading a One-Time Cost," the main implication is that we can shift inference costs from being tied to total text processed to being tied only to new information read. This means repeated work becomes nearly free, provided the data matches what’s already stored in Taliesin and Blaise.

Jane: That’s right, Tom. The paper suggests that for any workload that returns to the same documents, we should structure our systems so that the cost scales with how much *new* text is introduced, not by how much total text was ever processed.

Lu: I think the biggest potential here is enabling smaller models to effectively handle huge corpora because they don't need massive context windows if they can leverage this persistent memory layer to access only the most relevant parts.

Meng: We should keep an eye on how this stateful inference moves us toward building more cost-aware pipelines where the system dynamically chooses between loading saved memory or performing a full recomputation based on input similarity.

Lalam: For our team, this means we can build tools that are incredibly fast for routine tasks because the most common queries will be served instantly from the memory layer, freeing up resources for more complex reasoning.

Tom: That’s it for this session on Galahad; it really shows how precise memory management can fundamentally change how we think about scaling AI services. We’ll catch you next time as we look at something completely different in the research world.

Conclusion: Tom: So we’ve been diving deep into Galahad today, looking at how they managed to make reading text a one-time cost by using Taliesin and Blaise for byte-exact memory.

Jane: That’s right, Tom; essentially, they solved the problem of paying for repeated computation by treating stored data as a persistent asset rather than a fresh calculation every time.

Lu: I think the real power here is how that byte-exact matching works across different runtimes, which suggests we can achieve very stable state reuse even when moving between systems.

Meng: I see the practical value in that stability; if we can guarantee correctness through those checks, we get a much more predictable inference cost for our production pipelines.

Lalam: For us, this means our AI tools can become so much more responsive and reliable for our users because they won't be bogged down by repeated loading times when they ask the same questions.

Tom: Exactly; the paper really shows how precise memory management can fundamentally change how we think about scaling AI services.

Jane: It’s a fascinating look at moving inference from a stateless computation model to one with persistent memory, which is exactly what we need for many complex tasks.

Lu: We should definitely keep an eye on how this stateful approach interacts with other multimodal research, maybe seeing how it handles the complexity of visual or textual inputs in future work.

Meng: I’m curious if this memory persistence can be leveraged to reduce the sheer size of the models we have to deploy on edge devices.

Lalam: I hope we see this kind of efficiency bleeding into our core applications, making them feel much more intelligent and helpful every single time a user interacts with them.

Tom: Well, that wraps up our deep dive into "Working Around the Compute Ceiling: Byte-Exact Memory in Galahad Makes LLM Reading a One-Time Cost." We’ve seen how memory can be structured to reduce the cost of reading text significantly.

Jane: It was certainly an interesting look at how to manage computational budgets by focusing on reusing existing work.

Lu: Next week, we’re looking at something totally different, so get ready for a shift in focus as we explore new areas in multimodal reasoning and image-text alignment.

More episodes

← Home