Co-LMLM: Continuous-Query Limited Memory Language Models
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Co-LMLM: Continuous-Query Limited Memory Language Models".
Jane: Limited memory language models (LMLMs) externalize factual knowledge during pre-training to a knowledge base (KB), rather than memorizing it in their weights, and this work introduces Continuous-Query LMLM (CO-LMLM),
Tom: First, who's behind it and why it matters.
Title and authors: Tom: So we're diving into the paper "Co-LMLM: Continuous-Query Limited Memory Language Models," which tackles how language models handle facts by moving them out of their weights and into a knowledge base. Jane, can you give us the simple breakdown of what this paper is proposing?
Jane: Absolutely, Tom. Basically, this work introduces Continuous-Query LMLM or CO-LMLM, which takes the idea of externalizing knowledge during training but adds a flexible way to query that knowledge at inference time. Instead of just relying on a static key-value store with explicit queries, CO-LMLM pairs continuous keys with textual values, allowing it to generate vector queries in a minimal cost way while still pulling in human-readable and attributable facts for the final response.
Lu: That’s really interesting because the paper shows how this shifts the paradigm away from relational knowledge bases that were limiting the scaling of pre-training, especially when dealing with Wikipedia data where every article is centered on one entity <ref:2607.07707#pg1>.
Meng: From an engineering standpoint, moving to continuous vectors for queries seems efficient because it leverages the model's hidden state directly during generation rather than forcing it to decode a full natural language query at inference time.
Lalam: I see the implication here is that we gain a much more controllable system because the external memory isn't just static; it’s editable and supports direct unlearning through database operations, which is a big deal for maintaining factual accuracy over time.
Tom: So, moving beyond simple relational queries to continuous ones means the model can access a much wider range of knowledge without needing perfectly structured questions beforehand, which is really exciting for the flexibility of AI.
Jane: Exactly. They designed it to integrate this retrieval step seamlessly back into the autoregressive decoding process so that factual information becomes part of the generation itself, not just an appended afterthought.
Lu: The training methodology they used is clever because they jointly optimized next-token prediction with bidirectional contrastive losses using data pre-processed to simulate those KB queries, which really shows how integrated this system is from the start <ref:2607.07707#pg2>.
Title and authors: Meng: And that annotation pipeline sounds robust; moving from just Wikipedia to a scalable pipeline involving seed sets and fine-tuned annotators for spans suggests they’ve solved a major bottleneck in getting high-quality, tagged factual data at scale.
Lalam: That scalability is crucial because it means the knowledge base can grow organically with real-world text, not just being limited by whatever corpus we used for initial training.
Tom: Moving on to what this paper actually suggests they improved upon previous methods, what are the key advancements in CO-LMLM compared to earlier LMLM approaches?
Jane: The improvements focus heavily on the query mechanism and the resulting knowledge control. Specifically, they move away from relational KB structures toward continuous keys and values, which allows for flexible vector queries at minimal cost <ref:2607.07707#pg1>.
Lu: They’ve also addressed a major constraint mentioned in other works where removing or modifying specific facts usually requires additional training or specialized objectives, and CO-LMLM externalizes knowledge to make that modification possible through database operations instead of retraining <ref:2607.07707#pg2>.
Meng: That editability is what really interests me practically; if we need to update a piece of information, we can just adjust the external store rather than having to re-train the entire model weights, which simplifies maintenance immensely.
Lalam: That's fantastic for long-term system health because it decouples knowledge management from model parameter updates, which is a huge win for operational deployment.
Tom: And the performance results are quite compelling; they show that CO-LMLM outperforms prior LMLMs and vanilla LLMs across perplexity and factual precision benchmarks, even at smaller scales like three hundred sixty million parameters <ref:2607.07707#pg4>.
Jane: It’s not just about being better; they showed it can achieve factual precision scores that are in line with very advanced models like gpt-4o-mini and even surpass Claude Sonnet four point five, which is significant given the model size <ref:2607.07707#pg4>.
Lu: The evaluation across Simple English Wikipedia articles and general NLU scores confirms that the continuous query approach yields strong results when balanced with their contrastive retrieval loss <ref:2607.07707#pg3>.
Meng: While the performance is good, I wonder about the limitations they acknowledge; they mention that removing or modifying specific facts typically requires additional training or specialized objectives, which CO-LMLM tackles by externalizing knowledge to make that modification possible <ref:2607.07707#pg2>.
Title and authors: Lalam: Yes, and another limitation is that while it improves perplexity, the authors note that the scaling of pre-training still relies on a specific corpus like Wikipedia for their initial proof-of-concept demonstration <ref:2607.07707#pg1>.
Tom: So, to wrap up this discussion on Co-LMLM, we see a model that externalizes knowledge using continuous queries, which leads to better factual precision and provides a path toward editable and controllable memory systems. It's clear this paper moves the focus from just memorizing facts within weights to managing an external knowledge base dynamically.
Jane: That’s the core message: flexibility in querying without sacrificing grounding in verifiable facts through this continuous-query LMLM approach <ref:2607.07707#pg4>.
Lu: It opens up avenues for exploring how these external, editable stores can be leveraged with future agentic frameworks to create highly customized and persistent AI systems.
Meng: From an engineering view, the efficiency gain in forming the query—doing it in a single hidden-state forward pass rather than relying on explicit natural language queries—makes this much faster for real-time applications.
Lalam: Ultimately, this work confirms that externalizing knowledge simplifies unlearning and enhances controllability beyond what conventional LLMs offer <ref:2607.07707#pg2>.
Tom: So, to summarize the impact of Co-LMLM on the world right now, it’s about building AI systems that are not just big on parameters but are also smart managers of their external facts and can adapt those facts efficiently. We need to keep watching how this continuous query approach evolves.
Jane: It certainly gives us a solid direction for thinking about how we ground large language models more reliably in verifiable information, which is essential as AI becomes more integrated into critical decision-making processes.
Lu: The creative possibilities here are huge; imagine connecting these continuous vectors to complex, multi-modal knowledge graphs that the model can traverse dynamically <ref:2607.07707#pg1>.
Meng: I'm looking forward to seeing how this architecture translates into practical deployment scenarios where we need a model that can be quickly updated with new domain-specific facts without downtime.
Lalam: For me, the implication is that we can build AI cultures where knowledge isn't locked in; it’s fluid and manageable by the system itself, which feels like a step toward more truly intelligent agents.
The paper's summary: Tom: So we're talking about Co-LMLM, and what you've just explained is that this AI architecture takes factual knowledge out of the model's static weights and puts it into an external database that it can query continuously during generation.
Jane: Exactly, Tom, think of it like giving the AI a constantly accessible library instead of just trying to memorize every book inside its brain when you need an answer. This system lets the AI pull in relevant information from this external knowledge base based on the flow of its own thoughts in real-time.
Lu: It's wild because they're using continuous keys, not just fixed labels, to build these vectors, which means the AI can ask incredibly flexible questions that go beyond what a traditional database structure would allow <ref:2607.07707#pg1>. That opens up possibilities for knowledge traversal that we hadn't fully considered before.
Meng: From an engineering standpoint, I'm looking at how they handle the query formation itself; they generate these vectors from the model's hidden state at specific tokens, which is a much leaner way to get information than having the AI construct and process a full natural language search query every single time it needs a fact.
Lalam: And that lean mechanism is really powerful because it directly connects the model's internal reasoning path to external, verifiable data, ensuring that what the AI generates is grounded in something concrete <ref:2607.07707#pg4>. This level of direct integration into the generation process feels like a big step toward making AI outputs more trustworthy and attributable.
Tom: I'm really excited about that trust factor; if we can ensure the AI always cites its source from this external KB, it changes how we deploy these systems in sensitive areas. Jane, you mentioned the human-readable aspect—how does that actually help the end user?
Jane: It makes the facts accessible and understandable to people who aren't deep technical experts; instead of just seeing a raw retrieved string, they get a snippet that is clearly tied to a specific piece of knowledge, which builds confidence in what the AI is saying.
Lu: The implication for creative applications is huge; imagine an AI writing a story and being able to instantly check the veracity of every historical detail mentioned against this rich, external corpus without needing a separate search engine integration. That's a whole new level of context awareness <ref:2607.07707#pg1>.
Meng: Practically speaking, this means we can deploy these models in environments where the knowledge base is constantly updated—say, in fast-moving news or technical fields—without needing to halt the entire model for a full retraining cycle.
Lalam: That editability is what truly excites me; it means we're not locked into a static version of truth, which fundamentally alters how we think about maintaining an AI system over time. It moves us toward systems that are inherently adaptive and self-correcting in their knowledge base.
Tom: So, to recap, Co-LMLM provides a flexible way for an AI to pull in verifiable facts from a continuously queried external knowledge base, which boosts its factual accuracy and gives us powerful new ways to control what information it uses. Now that we've seen the summary, let's get into how they actually built this system during their training phase and how they managed to teach the model this retrieval skill in the first place.
The paper's improvements: Tom: So we’re getting into the specific advantages Co-LMLM brings to the table beyond just having an external knowledge base, and what Jane means is that they’ve fundamentally improved how the model interacts with that memory during its generation process.
Jane: Right, Tom, it's about making sure that when the AI needs a fact, it can fetch it dynamically in a single step without adding unnecessary delay or complexity to the overall process. They’ve focused on reducing the computational cost of retrieving knowledge as much as possible.
Lu: The move to continuous queries is key here because it means the model doesn't have to guess a perfect search term beforehand; it just uses its internal state at a specific point as a vector, which is way more fluid and adaptable than relying on fixed, pre-defined query structures.
Meng: That fluidity translates into real efficiency gains for deployment because forming that query takes less computational overhead per step than the methods we've seen before, like those involving explicit natural language queries at inference time.
Lalam: And this efficiency is what allows for truly dynamic and responsive AI systems; it means the AI can react to novel information with minimal latency, which is crucial for any real-time application that needs to keep up with a fast-moving environment.
Tom: It sounds like they’ve managed to solve the old problem where retrieval was a separate, heavy step tacked onto the main generation process; they’ve made it integrated into the core decoding loop. Jane, how does this integration actually make the final output better than just appending text afterwards?
Jane: When that retrieved snippet is spliced directly back into the context before decoding continues autoregressively, it means the fact becomes part of the sentence structure itself, which leads to more natural and coherent generation rather than just a tacked-on answer at the end.
Lu: This isn't just about adding text; it’s about creating a synergistic flow where retrieval informs every subsequent token prediction, leading to contextually richer and more grounded output across the whole response.
Meng: For practical implementation, this means we can build agents that don't just answer questions but can actively reference and synthesize facts from their memory in a way that feels like genuine reasoning rather than just pattern matching.
Lalam: This level of synergy is what I see as the most impactful vision; it moves AI from being a glorified search engine to being an active reasoner that constantly validates its claims against its internal, editable knowledge. It’s about creating an AI culture where information is inherently verifiable and integrated into every thought process.
Tom: So we’re talking about enhanced integration and efficiency, making the retrieval part of the thinking process itself, which really elevates the quality of the output substantially. Jane, what are you most impressed by regarding these specific improvements?
Jane: I'm really struck by how they maintain those crucial controllability features we saw in earlier LMLMs; even with this flexible query system, they kept the ability to edit and unlearn specific facts directly from the external database.
Lu: That is significant because it means we get the benefits of flexible querying without losing that essential layer of human-level control over the knowledge itself, which was a major sticking point in previous memory architectures.
Meng: I'm more focused on how this affects our roadmap; if we can implement this kind of targeted unlearning through database operations instead of full model retraining, our maintenance cycles for large language models could shrink dramatically.
Lalam: That capability to precisely remove specific knowledge entries without disrupting the entire model architecture is a massive win for long-term system reliability and ethical deployment. It gives us fine-grained control over the AI's worldview.
Tom: To wrap up this discussion on their improvements, it seems Co-LMLM successfully marries flexible, continuous querying with high fidelity fact retrieval while preserving the essential editability we need for reliable systems. Jane, you’ve laid out a clear path for how this architecture improves output quality and control in a very practical way. Now that we've covered the core benefits, let's take a moment to look at some of the specific experimental results they published on these benchmarks.
Conclusion: Tom: So we've covered the technical details of Co-LMLM, and what we’ve seen is that this work shows how integrating continuous retrieval directly into the generation process significantly boosts factual precision and control over knowledge management.
Jane: Exactly, Tom; it really demonstrates a more sophisticated way for AI to handle external information than just having a static memory bank, because it learns to use that memory in real time during its reasoning.
Lu: The creative potential here is immense; I’m thinking about using these continuous vectors not just for retrieval, but as parameters in complex generative processes that require constantly updating their factual understanding based on new input.
Meng: From an engineering viewpoint, the fact that they optimized the query formation cost so low means we can deploy these systems in high-throughput environments where real-time knowledge fetching is a genuine requirement.
Lalam: This work pushes us toward a future where AI isn't just a knowledge repository but an actively managed entity capable of dynamic adaptation and self-correction based on its continuously updated external facts, which fundamentally changes how we structure our thinking about AI systems.
Tom: It’s clear that the authors have built something very solid by successfully pairing continuous queries with their external memory system in the Co-LMLM architecture. Jane, what's your final thought on the overall impact this paper might have on how we design future language models?
Jane: I think it shows a really promising direction for grounding AI outputs in verifiable information, moving beyond just statistical fluency toward actual factual reliability and accountability.
Lu: It opens up exciting avenues for connecting these external, editable stores to complex, multi-modal knowledge graphs that the model can traverse dynamically in ways we haven't even fully conceptualized yet.
Meng: For us on the engineering side, the ability to perform targeted unlearning through database operations instead of full retraining is a major operational advantage that will simplify how we manage and update these models in production.
Lalam: I see this as a huge step for AI culture; it suggests a move toward systems that are inherently adaptive and trustworthy because their knowledge isn't locked down, but fluid and actively managed by the system itself.
Tom: We’ve really seen how Co-LMLM uses continuous queries to create more flexible, efficient, and controllable memory systems for language models. It’s a fascinating piece of research that shows the power of externalizing knowledge in a sophisticated way.
Jane: It gives us a really clear blueprint for building AI that can be both creative and rigorously grounded in verifiable data simultaneously.
Lu: We're going to be looking closely at how this continuous vector space interacts with other memory systems we've been developing, like those focusing on long-form video state management.
Meng: I want to see if the efficiency gains translate into practical, low-latency applications where the retrieval cost is truly minimal in a live setting.
Lalam: Ultimately, this paper confirms that externalizing knowledge simplifies unlearning and enhances controllability far beyond what conventional LLMs offer.
Tom: That’s our wrap-up on Co-LMLM; it’s a really exciting development in making AI smarter about managing its own facts. Next time, we'll be looking at how this concept applies to agentic workflows and autonomous decision-making.
Cornell University
cs.CL, cs.AI, cs.LG
Submitted: 2026-07-08
Updated: 2026-10-02
Code: https://github.com/shmsw25/FActScore
Project page: https://lil-lab.github.io/co-lmlm-web
Importance score: 83/100
The gist: Limited memory language models (LMLMs) externalize factual knowledge during pre-training to a knowledge base (KB), rather than memorizing it in their weights, and this work introduces
Key concepts
- Continuous-Query LMLM (CO-LMLM)
- A decoder-only Transformer designed to act as both a language model and a dense retriever. It generates query vectors similar to tool calling, using the hidden state at a special token to perform retrieval against an external key-value store, then splicing the retrieved fact into the context before continuing generation.
- Knowledge Base (KB) Structure
- The KB in CO-LMLM stores 'vector keys and string values' instead of traditional relational tuples. The vector keys are generated by the model, and their corresponding values are factual text spans extracted from pre-training data. This structure allows for flexible querying without relying on fixed relational schemas.
- Unlearning Capability
- Externalizing knowledge simplifies unlearning; instead of retraining the entire model, relevant entries can be removed directly from memory. CO-LMLM successfully retains knowledge outside the forget set, unlike other methods where entanglement degrades retention performance.
Terminology
Summary
Limited memory language models (LMLMs) externalize factual knowledge during pre-training to a knowledge base (KB), rather than memorizing it in their weights, and this work introduces Continuous-Query LMLM (CO-LMLM), which pairs continuous keys with textual knowledge values to allow for flexible vector queries at minimal cost while integrating human-readable and attributable retrieved knowledge into generation. This paradigm offers multiple advantages, including knowledge control capabilities that remain beyond conventional LLMs.
How it works
The CO-LMLM model is a decoder-only Transformer designed to function as both a language model and a dense retriever. It operates by generating query vectors similar to tool calling, where the retrieval query is not represented as text but as the model’s hidden state at a special token. During inference, when the model emits the special token, it reads the last-layer hidden state at that position to perform a retrieval operation against an external key-value store and retrieve a top-1 fact text snippet. This retrieved snippet is then spliced into the context before decoding resumes autoregressively.
Training Methodology
The training involves jointly optimizing next-token prediction (NTP) and bidirectional contrastive losses using pre-training data that has been pre-processed to simulate KB queries. The training data consists of raw text where factual spans are enclosed by and tags, each paired with a natural-language question whose answer is the span. The model minimizes the NTP loss, which is formulated as LNTP(θ) = −Pt∈M / log p(xt x and. Simultaneously, a contrastive retrieval loss (LCL) is minimized over synthesized positive and negative pairs. The positive pairs consist of the retrieval vectors at the query markers appended to questions, while negative pairs consist of all other questions.
Knowledge Base Construction and Indexing
The KB in CO-LMLM stores vector keys and string values instead of relational tuples.
These vector keys are produced by the model, and their corresponding values are factual text spans extracted from the pre-training data. A dense retrieval index is built from this pre-training data by running the pre-trained language model over the corpus to extract L2-normalized hidden states at every position, storing these states as keys in a dense vector index whose values are the verbatim in-text span at that position.
Data Annotation Pipeline
A scalable annotation pipeline is employed to tag free-form factual spans in arbitrary text, moving beyond prior restrictions to Wikipedia. This pipeline involves three stages:
-
Annotate a
seed set
using a state-of-the-art frontier LLM (like Gemini 3.1 Pro) with the format span. -
Fine-tune lightweight annotators on this seed data: a fact span annotator (a ModernBERT encoder fine-tuned for BIO token classification) and a question generator (a Qwen2.5-1.5B-Instruct decoder fine-tuned with LoRA).
-
Apply these annotators to the full pre-training corpus, converting raw documents into the annotated pre-training corpus for CO-LMLM.
Evaluation and Performance
CO-LMLM is evaluated across several benchmarks, including perplexity on Simple English Wikipedia articles, factual precision (FactScore), and general NLU scores. Across model scales, CO-LMLM outperforms prior LMLMs and vanilla LLMs in both perplexity and factual precision. For instance, at the 360M scale, it achieves lower perplexities than models pretrained on 40× more data (like SMOLLM2-360M) and SimpleQA-verified performance that is in line with gpt-4o-mini and higher than Claude Sonnet 4.5.
Furthermore, CO-LMLM retains the controllability benefits of LMLM, as its external memory remains editable and supports direct unlearning through database operations.
Unlearning Capability
Externalizing knowledge simplifies unlearning: instead of retraining the model, one simply removes relevant entries from memory. The study tests whether CO-LMLM preserves this benefit while replacing relational queries with continuous queries. It demonstrates that CO-LMLM retains knowledge outside the forget set, unlike other methods that degrade retain-set performance because of the entanglement of knowledge storage,
confirming its controllability and editability.
Inference Efficiency
CO-LMLM is shown to be more efficient in terms of query formation cost compared to prior methods. For instance, CO-LMLM forms its query in a single " 2.
Improvements for AI systems
Here are specific improvements to AI systems based on the CO-LMLM (Continuous-Query Limited Memory Language Models) paper:
-
The AI system can externalize factual knowledge into a scalable, editable, and human-readable Knowledge Base (KB) during pre-training, rather than relying solely on static model weights.
-
The system can perform continuous retrieval during generation by treating the model's hidden state at specific tokens as a flexible query vector. This allows for dynamic, context-aware knowledge fetching without the overhead of decoding explicit natural language queries at inference time.
-
The system can generate flexible, free-form vector queries (continuous vectors) that span a much broader set of knowledge than models constrained by relational KB structures (like REL-LMLM).
-
The system can integrate human-readable and attributable retrieved knowledge directly into its generation process, ensuring outputs are grounded in verifiable facts.
-
The system gains superior factual precision and performance on short-form QA tasks (SimpleQA), long-form factual generation (FactScore), and complex reasoning tasks compared to vanilla LLMs and prior LMLM variants, even at smaller model scales (e.g., 360M scale).
-
The system supports efficient knowledge unlearning through direct database operations on the external KB, allowing for targeted removal of specific facts without requiring expensive full model retraining or parameter updates (Contrast with NPO).
-
The system can operate efficiently in a dynamic setting where retrieval is triggered only when necessary, and its per-retrieval query formation cost is significantly lower than models that generate explicit natural language queries (e.g., CO-LMLM's single hidden-state forward vs. LMLM-ASKER's 28ms).
-
The system can scale knowledge acquisition beyond specific corpora (like Wikipedia) by training on diverse web data (FineWeb-Edu), demonstrating robustness and improved performance in factual recall on unseen domains.
-
The system can be tuned to prioritize high information density facts, ensuring that retrieved knowledge is meaningful and relevant, avoiding the tagging of common knowledge or incidental details.
Sources
- jina-embeddings-v5-text: Task-Targeted Embedding Distillation
- Memory Layers at Scale
- PIQA: Reasoning about Physical Commonsense in Natural Language
- Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- The Faiss library
- Toy Models of Superposition
- SimpleQA Verified: A Reliable Factuality Benchmark to Measure Parametric Knowledge
- Atlas: Few-shot Learning with Retrieval Augmented Language Models
- MemLLM: Finetuning LLMs to Use An Explicit Read-Write Memory
- Generative Representational Instruction Tuning
- Learn to Memorize: Scalable Continual Learning in Semiparametric Models with Mixture-of-Neighbors Induction Memory
- Pretraining with hierarchical memories: separating long-tail and common knowledge
- Qwen2.5 Technical Report
- Representation Learning with Contrastive Predictive Coding
- Measuring short-form factuality in large language models
- $\text{Memory}^3$: Language Modeling with Explicit Memory
- ImpRAG: Retrieval-Augmented Generation with Implicit Queries
- Pre-training Limited Memory Language Models with Internal and External Knowledge
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering