Training Large Language Models to Reason in a Continuous Latent Space
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Training Large Language Models to Reason in a Continuous Latent Space".
Jane: The paper was written by Shibo Hao, Sainbayar Sukhbaatar, DiJia Su, Xian Li, Zhiting Hu et al. from Meta and University of California San Diego.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Improvements: Tom: Now that we understand the core idea, let’s look at the actual results from "Training Large Language Models to Reason in a Continuous Latent Space." The authors claim that this new approach leads to some very impressive performance improvements on real-world problems.
Jane: For instance, on the GSM8k math reasoning dataset, Coconut achieves significantly higher accuracy compared to the standard CoT approach, which is a huge win for complex problem-solving.
Lu: And I’m particularly interested in their findings regarding logical reasoning datasets like ProsQA; it seems like that continuous space is much better equipped to handle problems that require extensive planning and backtracking.
Meng: The data suggests a clear trade-off improvement as well; you get better accuracy without having to generate as many tokens, which translates directly into faster inference times for production systems.
Lalam: This efficiency gain means we can achieve deeper reasoning in the AI without multiplying the computational cost of every single word that gets generated, which is a massive win for resource management.
Tom: The paper’s explanation of why this works is that by allowing for this latent search, the model avoids those early mistakes and hallucinations that are common when it commits to a wrong path in language space.
Jane: It's about having the freedom to explore; even if they start with an incorrect idea, they can backtrack and refine their continuous thought until we find the correct answer through iterative refinement.
Lu: The fact that this emergent BFS behavior happens without explicitly training it is what I find most surprising, it suggests a natural optimization of the architecture itself.
Meng: From an engineering standpoint, having a mechanism that inherently avoids premature commitment would drastically reduce error rates in mission-critical applications where accuracy is paramount.
Lalam: It allows for a more robust and reliable intelligence; rather than just guessing based on language patterns, we are enabling actual systematic exploration of possibilities.
Tom: The evidence shows that this continuous search is not only powerful but also very efficient, which is something the researchers were excited to demonstrate with their findings on "Training Large Language Models to Reason in a Continuous Latent Space."
Jane: It proves that we can achieve high-level reasoning by finding multiple paths simultaneously rather than just one single path.
Lu: The emergence of this broad search strategy suggests the model is optimizing its internal representation to find the most robust solution.
Meng: It's a practical advantage, meaning we can deploy systems that perform complex logical tasks with much higher confidence and lower latency.
Lalam: This ability to explore multiple paths allows us to build AI that truly understands complexity, moving beyond mere pattern matching.
Conclusion: Tom: We've spent a lot of time discussing the methodology and the impressive results of "Training Large Language Models to Reason in a Continuous Latent Space," so how does this fundamentally change things for us moving forward?
Jane: It proves that we might not need to force all our AI reasoning through language, which opens up possibilities for much more complex tasks than we thought possible before.
Lu: I think the biggest impact is that this opens up the possibility of training LLMs in a way that truly mirrors how human cognition operates, allowing us to build machines capable of genuine exploration.
Meng: We need to look at how we can integrate these continuous thoughts into our existing pipelines, making sure we can utilize both this efficiency and robustness in real-world applications.
Lalam: My vision is that this will accelerate the path toward an AI that doesn't just generate plausible text, but one that has a genuinely sophisticated internal model of the world.
Tom: We have to acknowledge the work of all these researchers who proved that language space is not always optimal for reasoning and "Training Large Language Models to Reason in a Continuous Latent Space" is really paving the way for much better AI.
Jane: It’s a bold step toward letting our machines think without linguistic constraints, which is something I think we should all be excited about as it was long overdue.
Lu: I'm excited to see what further research into this latent space yields, especially when combining this idea with other advanced techniques like superposition for solving complex problems.
Meng: We're looking forward to scaling up this technology and making sure the engineering challenges are addressed with the insights gained from these future steps.
Lalam: It’s a fundamental shift that promises a more capable and reliable form of intelligent life for all stakeholders in culture.
Tom: This is such an exciting area, Jane, understanding how we transition from language to this continuous thought space within "Training Large Language Models to Reason in a Continuous Latent Space."
Jane: We are definitely looking forward to see how much further this goes after "Training Large Language Models to Reason in a Continuous Latent Space" is presented.
Lu: I can already imagine the new capabilities this opens up for complex problem solving that was previously out of reach for us.
Meng: It's going to make designing these scalable AI systems so much more feasible, which will be a huge relief for the engineering teams.
Lalam: This paper offers a glimpse into a future where our AI is truly capable of independent, deep thought and reasoning about the world.
Paper discussion segment 3: Tom: To wrap up our deep dive into "Training Large Language Models to Reason in a Continuous Latent Space," the key takeaway is that we are fundamentally shifting from predicting the next word to modeling the entire internal state of thought.
Jane: And what that means for system design is massive. Previously, if an LLM got stuck on a wrong premise—say, it misinterpreted a unit of measurement or misunderstood a causal link—it had no natural mechanism to pause and correct that core belief without explicit prompt engineering telling it *how* to backtrack.
Lu: But by operating in this continuous latent space, the model isn't forced into that linear, textual commitment. It’s essentially keeping an active, multi-dimensional ledger of all possible interpretations simultaneously. This allows for a form of internal self-correction that is far more robust than any prompting technique we currently use.
Meng: From an engineering standpoint, this suggests a modularity we haven't been able to achieve before. We can begin to design AI agents where the "reasoning engine" operates independently of the "language output layer." The engine processes complex logic—like simulating physics or running a detailed multi-step plan—and only translates its final, verified continuous state into natural language when it’s absolutely ready.
Lalam: That separation is crucial because it acknowledges that thinking is not the same thing as speaking. Our current models are forced to perform both tasks simultaneously, which inherently limits the depth of their internal reasoning capacity. This method gives us a clean pathway to decoupling those two processes for specialized applications.
Tom: So, if I understand correctly, we aren't just making smarter text generators; we are building systems with genuinely independent cognitive components—a sophisticated "thought module" that feeds into a communicative "expression module."
Jane: Exactly. And this capability opens the door to entirely new classes of AI tools—things like advanced scientific discovery assistants or complex strategic planners—where the cost of an early mistake is prohibitively high, and we need guaranteed reliability.
Lu: It moves us from a system that *sounds* intelligent to one that *is* systematically capable of deep, verifiable reasoning. This changes the entire benchmark for what we consider "intelligent" in artificial systems.
Meng: It means our focus shifts from optimizing token generation rates to optimizing the fidelity and complexity of the latent state itself. That's a much more challenging, but ultimately more rewarding, engineering goal.
Lalam: And this brings us to where we are now: understanding that the future of AI isn't just about bigger models; it’s about fundamentally different architectures that allow for true cognitive freedom. Now that we understand the potential of continuous thought, I think we need to explore how these systems can interact with each other.
Conclusion: Tom: We've covered so much ground today, from the mechanics of Coconut to its incredible results in "Training Large Language Models to Reason in a Continuous Latent Space." The consensus is that we are moving toward a point where AI systems can truly think beyond human language constraints.
Jane: That’s the most encouraging part for me; knowing that our AI has the potential to develop its own internal, verifiable thought processes without needing external instruction is a massive leap forward for building reliable systems.
Lu: I agree, and from a theoretical standpoint, it feels like we've cracked a major constraint in the architecture of how we view LLMs. This shift from discrete language to continuous latent space really opens up entirely new ways of thinking about cognitive modeling.
Meng: It's vital that we take these findings seriously when it’s time to start building production systems, ensuring that this methodology is robust enough to handle real-world scale and complexity.
Lalam: This move toward independent, continuous thought is a profound change for us. It allows the AI to become a more sophisticated partner in discovery, contributing deeper insights into the fabric of human culture and knowledge.
Tom: Indeed, it’s clear that "Training Large Language Models to Reason in a Continuous Latent Space" is suggesting that we are no longer limited by what's on the surface of language.
Jane: We've seen evidence that this continuous reasoning not only increases accuracy but also manages the efficiency of token usage beautifully.
Lu: I think the real excitement lies in seeing how this will interact with other advanced techniques like superposition, leading to even more powerful systems down the road.
Meng: We're eager to see how these insights translate into concrete, scalable engineering solutions that make sense for large-scale deployment.
Lalam: It promises a future where our AI is truly capable of independent, deep thought about its surroundings and society.
Shibo Hao, Sainbayar Sukhbaatar, DiJia Su, Xian Li, Zhiting Hu, Jason Weston, Yuandong Tian
Meta · University of California San Diego
cs.CL
Submitted: 2026-08-23
Updated: 2026-08-25
Comments: Accepted to COLM 2025
Code: https://github.com/facebookresearch/coconut
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 76/100
The gist: The paper, "Training Large Language Models to Reason in a Continuous Latent Space," introduces a novel paradigm called Coconut (Chain of Continuous Thought) designed to overcome the inherent
Key concepts
- Continuous Latent Space
- This refers to a model's internal representation of thought, moving beyond simple language patterns. Instead of committing to a single textual path, the the model maintains multiple interpretations simultaneously. This allows it to explore various possibilities and refine its understanding until it finds the correct answer.
- CoT (Chain-of Thought) Approach
- This is a standard method where LLMs generate reasoning by breaking down problems into sequential steps using language. The new continuous latent space approach significantly outperforms this traditional method on complex datasets, achieving higher accuracy and better handling of planning and backtracking.
- Backtracking/Refinement
- The continuous latent space allows the model to explore an initial incorrect idea without committing to it. It can then refine or backtrack from that starting point by iterating through multiple interpretations simultaneously, ensuring a more robust path toward the correct solution.
Terminology
Summary
The paper, Training Large Language Models to Reason in a Continuous Latent Space,
introduces a novel paradigm called Coconut (Chain of Continuous Thought) designed to overcome the inherent limitations of using natural language for complex reasoning in Large Language Models (LLMs).
Problem Statement and Motivation
The authors argue that current LLMs are restricted to reason within the language space,
where they typically express their process using Chain-of-Thought (CoT). However, this approach is suboptimal. The paper states: the amount of reasoning required for each particular token varies greatly, yet current LLM architectures allocate nearly the same computing budget for predicting every token.
Furthermore, while some tokens are generated solely for fluency,
others require complex planning and pose huge challenges to LLMs.
The authors propose that it would be ideal for LLMs to have the freedom to reason without any language constraints
before translating findings into language.
The Coconut Paradigm
To explore this potential, the authors introduce Coconut. This involves a modification where the model switches between language mode
and latent mode. In latent mode, it directly utilizes the last hidden state as the next input embedding.
This last hidden state is termed a “continuous thought.” The core mechanism is that instead of decoding this into a word token, we feed it back to the LLM as the subsequent input embedding directly in the continuous space.
Mechanism and Training Procedure
The training process leverages existing language CoT data through a multi-stage curriculum inspired by Deng et al. (2024). In this procedure, at each stage, we integrate c additional continuous thoughts... and remove one language reasoning step.
The model is trained to optimize the normal negative log-likelihood loss while masking the loss on questions and latent thoughts. Because these continuous thoughts are fully differentiable, the system can be optimized end-to-end by gradient descent.
Theoretical Advantage: Latent Tree Search
The authors demonstrate that this shift to a latent space enables an advanced reasoning pattern. Unlike language-based reasoning, continuous thoughts in Coconut can encode multiple alternative next reasoning steps, allowing for a reasoning process akin to breadth-first search (BFS) to solve the problem.
This allows the model to perform a breadth-first search (BFS) to solve the problem, rather than prematurely committing to a single deterministic path like CoT.
The latent space facilitates this tree search. The authors observe that the continuous representation of reasoning can encode multiple alternative next reasoning steps,
which allows the model to progressively eliminate incorrect paths and converge on the correct answer.
Experimental Results
The performance of Coconut is evaluated across several datasets:
-
Logical Reasoning (ProsQA): On ProsQA, a dataset requiring
stronger planning ability,
Coconut and some of its variants even surpasses language-based CoT methods, while generating significantly fewer tokens during inference.
-
Math Reasoning (GSM8k): For math reasoning,
using continuous thoughts is shown to be beneficial to reasoning accuracy, mirroring the effects of language reasoning chains.
-
Efficiency: Table 1 shows that Coconut achieves a
superior trade-off between reasoning efficiency and accuracy
compared to CoT.
The authors further analyze the advantage of latent space for planning tasks, noting that nodes with lower heights (closer to terminal states) consistently receive more accurate and definitive probability evaluations,
while nodes with greater heights exhibit more ambiguous evaluations.
This suggests that by delaying deterministic decisions, latent reasoning enhances the model’s ability to differentiate correct paths from incorrect ones.
Conclusion
The paper concludes that Coconut effectively enhances LLM performance across a variety of reasoning tasks
and demonstrates that "continuous thoughts can represent multiple alternative next steps. This enables the model to perform BFS over possible reasoning paths, rather than prematurely committing to a single deterministic trajectory as in language space CoT reasoning."
Improvements for AI systems
Based on the research presented in Training Large Language Models to Reason in a Continuous Latent Space, I have identified several critical, specific improvements that can be implemented into an existing LLM architecture. These changes move beyond simple prompt engineering and require fundamental modifications to both training and inference pipelines.
The Change:
Modify the standard autoregressive decoding loop during inference. Instead of relying on the softmax output of a language model head to select a next token (softmax(W h t)), we implement a Latent Mode where the LLM's final hidden state (h t) is directly used as the input embedding for the subsequent step.
-
Mechanism: For a sequence of tokens x = (x 1,, x T), when the model enters latent mode (marked by bot and eot), we replace the token embedding at position t+1 with the previous hidden state: E t+1 = [, e(x t), h t,].
-
Implementation Detail: This requires that the final normalization layer of h t remains stable (as noted in Section 3) and that the input embedding space is designed to accept these continuous vectors without requiring a mapping back to the discrete token vocabulary.
What this enables:
The LLM can now perform unrestricted latent reasoning, bypassing the constraints of natural language coherence. This allows the model to maintain abstract, multi-dimensional representations of intermediate variables and relationships, crucial for complex planning.
-
Process: The training data, containing explicit language reasoning steps (CoT), is systematically modified across stages. At stage k, we replace k sequential language reasoning steps with k times c continuous thoughts, where c is a hyperparameter controlling the parallelism (c=1 or 2).
-
Training Objective: The cross-entropy loss is calculated only on the remaining text tokens (the input/output sequence) after the continuous thoughts are masked.
-
Implementation Detail: The optimizer state must be reset whenever transitions between stages occur, ensuring the the model learns a distinct, non-greedy strategy at each stage.
-
Strategy: While we can use a binary classifier trained on the latent thoughts (Strategy A), a simpler, highly effective method is to enforce a fixed termination length or padding (Strategy B).
-
Mechanism: The model starts in Latent Mode (bot). Upon completion, it transitions back to Language Mode using an eot token. This ensures the final output is readable and grounded in natural language, while the internal process remains continuous.
The resulting improved AI system will possess the following capabilities:
-
Superior Planning: It can execute complex reasoning tasks (e.g., complex logical deduction, multi-step mathematical problems) that require a deep search strategy, not just a linear path.
-
Reduced Hallucination: By evaluating multiple paths simultaneously in latent space and avoids premature commitment to an incorrect branch, the system significantly reduces the rate of generating invalid or unsupported reasoning steps (e.g.,
hallucinating
edges). -
Enhanced Efficiency: For complex problems, it achieves higher accuracy while generating significantly fewer tokens compared to traditional CoT methods, reducing computational cost and inference latency.
-
Autonomous Complex Problem Solving: It can handle long-tail distribution of reasoning chains (e.g, 6+ steps) with high reliability, even when operating without explicit language guidance.
Sources
- GPT-4 Technical Report
- Large Concept Models: Language Modeling in a Sentence Representation Space
- Hopping Too Late: Exploring the Limitations of Large Language Models on Multi-Hop Queries
- Training Verifiers to Solve Math Word Problems
- Implicit Chain of Thought Reasoning via Knowledge Distillation
- From Explicit CoT to Implicit CoT: Learning to Internalize CoT Step by Step
- The Llama 3 Herd of Models
- Looped Transformers for Length Generalization
- Stream of Search (SoS): Learning to Search in Language
- Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach
- Energy-Based Transformers are Scalable Learners and Thinkers
- Think before you speak: Training Language Models With Pause Tokens
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Reasoning with Language Model is Planning with World Model
- LLM Reasoners: New Evaluation, Library, and Analysis of Step-by-Step Reasoning with Large Language Models
- Teaching Large Language Models to Reason with Reinforcement Learning
- Decomposed Prompting: A Modular Approach for Solving Complex Tasks
- Beyond A*: Better Planning with Transformers via Search Dynamics Bootstrapping
- Chain of Thought Empowers Transformers to Solve Inherently Serial Problems
- Text and Patterns: For Effective Chain of Thought, It Takes Two to Tango
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering