Latent Chain-of-Thought as Planning: Decoupling Reasoning from Verbalization
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "Latent Chain-of-Thought as Planning".
Tom: Chain-of-Thought (CoT) reasoning has revolutionized Large Language Models (LLMs) by decomposing complex problems into sequences of intermediate steps,
Jane: First, who's behind it and why it matters.
Title and authors: Tom: Moving on to a deeper look at the actual summary of "Latent Chain-of-Thought as Planning: Decoupling Reasoning from Verbalization." They explain that PLaT models reasoning as a deterministic trajectory of latent planning states rather than just relying on fixed steps in token sequences.
Jane: It really boils down to this: they’ve introduced the concept of a latent Planner that autoregressively evolves these states in a high-dimensional continuous manifold, while the Decoder is responsible for grounding those thoughts into text when necessary.
Lu: Essentially, they are proposing that the core reasoning process happens implicitly within this continuous latent space, and we only get verbalization when an external interface is required to communicate with the outside world.
Meng: I see them saying that this approach contrasts sharply with previous end-to-end implicit methods that relied on a fixed number of latent steps during inference, which is what we’ve been trying to move away from.
Lalam: They explicitly state that PLaT models reasoning as a deterministic trajectory of latent planning states while using a separate Decoder to ground these thoughts into text when necessary, which is the main contribution.
Tom: So, the major summary point is this decoupling: the Planner generates the thought process in latent space, and a separate Decoder handles translating those thoughts into natural language only when that translation is required.
Jane: And because of that structure, it allows for dynamic termination of reasoning based on semantic needs rather than relying on fixed hyperparameters during inference, which gives it much more flexibility.
Lu: That dynamic control over when to stop the reasoning process seems like a very important feature because it makes the system behave more like human System two thinking by only committing to language when necessary.
Meng: It shifts the burden from the model having to commit to a fixed path of token generation toward managing a trajectory in continuous latent space.
Lalam: They also detail how they use dedicated linear projectors to bridge the gap between the LLM backbone dimension and this latent dimension, which is essential for mapping history into that new space.
Tom: That projector setup is what makes the entire mechanism functional, allowing history to be mapped into that hidden state representation so the Planner can start evolving its trajectory.
Jane: And this whole process allows for a richer internal representation than just a fixed sequence of tokens, which is what we’ve been working on in terms of internal model capacity.
Lu: It’s fascinating how they manage the flow between these different spaces; it suggests that mapping high-dimensional thought onto a continuous latent space and back again is where the real potential lies for future research directions in understanding these internal dynamics.
Meng: For us, this means we can potentially deploy models that are more efficient because they aren't generating tokens sequentially through the whole reasoning process every single time.
Lalam: The paper’s main contribution is enabling dynamic termination of latent planning and intermediate interpretability of latent states, which we need to flag as the main things to remember from PLaT.
The paper's summary: Tom: Let's talk about the specific improvements they suggest, because PLaT isn't just a concept; it has concrete mechanisms like SFT and RL stages built into its training process to make this work.
Lu: They detail the training pipeline, starting with an SFT stage where the Planner autoregressively steps forward to generate latent states in response to a question, and then the Decoder uses those projected states as a prefix to verbalize them.
Meng: That initial reconstruction loss calculated as the crossentropy between ground-truth text and what the model predicts conditioned on state Sk is how they anchor those latent states in actual textual data, which is crucial for supervised learning.
Lalam: For robustness, they inject Gaussian Noise into accumulated states during training to improve the system’s ability to handle unexpected inputs during inference.
Tom: After that comes the RL stage where they use Reinforcement Learning to optimize only the Decoder parameters while freezing all Planner parameters, which is a smart way to keep the reasoning structure stable while refining how it speaks.
Jane: They employ a Decoupled GRPO objective for this refinement, which introduces diversity by enabling temperature sampling in the Decoder from fixed latent states Sk, so they aren't locked into just one output mode.
Lu: This is where they move toward policy refinement by optimizing the linguistic policy without disturbing the underlying reasoning logic during this crucial stage.
Meng: The reward function provides dense supervision based on semantic validity and mathematical correctness for both intermediate steps and the final answer, which gives a lot of feedback to the Decoder to focus on producing high-quality text.
Lalam: And they also include prompts designed to leverage LLMs like GPT-4o-mini for validation, where an automated evaluator verifies step validity with high consistency, which is used directly in the RL reward function.
Tom: It sounds like the final result is a system that’s trained to generate reasoning steps that are semantically valid and mathematically correct, and then the Decoder learns to translate those validated latent thoughts into language.
Jane: So they are essentially using these specific training techniques to ensure that the output is not just fluent text, but text grounded in verified latent planning.
Lu: The way they approach this reinforcement learning stage suggests that optimizing the linguistic policy without disturbing the underlying reasoning logic during this crucial stage is a sophisticated way to handle policy refinement.
Meng: It’s about seeing how this translates into tangible gains in efficiency and reliability on real-world applications we can actually measure.
The paper's improvements: Tom: So we’ve gone through the PLaT framework and its improvements, and I think we've covered the main points of this paper on "Latent Chain-of-Thought as Planning: Decoupling Reasoning from Verbalization." It really boils down to moving reasoning into a continuous latent space for planning.
Jane: The implications are that we can have more flexible systems that can dynamically decide when to stop thinking based on what they need, which is a major step toward building more flexible AI.
Lu: I think the way they approach the mapping between continuous space and discrete space is where the real potential lies for future research directions in understanding these internal dynamics.
Meng: Practically speaking, this means we can potentially deploy models that are more efficient because they aren't generating tokens sequentially through the whole reasoning process every single time.
Lalam: The paper’s main contribution is enabling dynamic termination of latent planning and intermediate interpretability of latent states, which we need to flag as the main things to remember from PLaT.
Tom: Exactly, so we wrap up our discussion on "Latent Chain-of-Thought as Planning: Decoupling Reasoning from Verbalization." It’s a really solid piece of work that gives us a clear path forward for building more sophisticated AI agents.
Jane: I think the idea of separating thought from language is something we should keep emphasizing in our discussions as we move forward.
Lu: And I think the way they approach the mapping between continuous space and discrete space is where the real potential lies for future research directions in understanding these internal dynamics.
Meng: And for us, it’s about seeing how this translates into tangible gains in efficiency and reliability on real-world applications we can actually measure.
Lalam: The paper's main contribution is enabling dynamic termination of latent planning and intermediate interpretability of latent states, which we need to flag as the main things to remember from PLaT.
Conclusion: Tom: So we’ve covered a lot about PLaT, and we've seen how they're taking reasoning into that continuous latent space to plan rather than just following a strict token sequence.
Jane: And the main takeaway for us is how this separation of thought from language opens up new flexibility for AI systems moving forward.
Lu: I think the way they manage that mapping between high-dimensional thought and continuous latent space is where it really opens up new avenues for interpreting those internal structures we’ve been studying.
Meng: Practically speaking, this means we can potentially deploy models that are more efficient because they aren't generating tokens sequentially through the whole reasoning process every single time.
Lalam: The paper’s main contribution is enabling dynamic termination of latent planning and intermediate interpretability of latent states, which we need to flag as the main things to remember from PLaT.
Tom: Exactly, so we wrap up our discussion on "Latent Chain-of-Thought as Planning: Decoupling Reasoning from Verbalization." It’s a really solid piece of work that gives us a clear path forward for building more sophisticated AI agents.
Jane: I think the idea of separating thought from language is something we should keep emphasizing in our discussions as we move forward.
Lu: And I think the way they approach the mapping between continuous space and discrete space is where it really opens up new avenues for interpreting those internal structures.
Meng: And for us, it’s about seeing how this translates into tangible gains in efficiency and reliability on real-world applications we can actually measure.
Lalam: The paper's main contribution is enabling dynamic termination of latent planning and intermediate interpretability of latent states, which we need to flag as the main things to remember from PLaT.
Beihang University
cs.AI, cs.CL
Submitted: 2026-01-29
Updated: 2026-09-28
Code: https://github.com/yunsaijc/PLaT
Importance score: 81/100
The gist: Chain-of-Thought (CoT) reasoning has revolutionized Large Language Models (LLMs) by decomposing complex problems into sequences of intermediate steps, but it suffers from computational cost and
Key concepts
- Latent Planner
- This module autoregressively evolves a sequence of planning states in a high-dimensional continuous space. It maintains a probabilistic density over various logical possibilities, effectively simulating the internal thought process as a trajectory rather than relying on fixed steps.
- Decoder for Verbalization
- The Decoder takes the evolved latent states from the Planner and aggregates them using an Exponential Moving Average (EMA) to create stable representations. It then projects these aggregated states into text, allowing reasoning to be grounded in language only when necessary.
- Dynamic Termination
- Instead of relying on pre-set limits, PLaT allows reasoning to stop dynamically. The system pauses generation if the Decoder's output does not match a specific termination signal ('tans'), enabling flexible and context-aware reasoning lengths.
- Decoupled GRPO
- Reinforcement Learning is used to refine only the Decoder parameters while keeping the Planner fixed. This objective introduces diversity by sampling from fixed latent states during training, encouraging the model to explore a broader solution space.
Terminology
Summary
Chain-of-Thought (CoT) reasoning has revolutionized Large Language Models (LLMs) by decomposing complex problems into sequences of intermediate steps, but it suffers from computational cost and reasoning path collapse due to its reliance on discrete token spaces. This paper introduces PLaT, a framework that reformulates latent reasoning as planning by fundamentally decoupling the core reasoning process from verbalization.
The gist
PLaT models reasoning as a deterministic trajectory of latent planning states while using a separate Decoder to ground these thoughts into text when necessary, allowing for dynamic termination of reasoning rather than relying on fixed hyperparameters.
How it works
The PLaT architecture comprises two distinct modules: a latent Planner and a Decoder for verbalization. The Planner autoregressively evolves a trajectory of states in a high-dimensional continuous manifold, maintaining a probabilistic density over multiple logical possibilities until a decision is required. This contrasts with previous end-to-end implicit methods that rely on fixed numbers of latent steps.
The architecture utilizes dedicated linear projectors to bridge the LLM backbone dimension and the latent dimension. The Planner initializes the trajectory using an encoder projector, maps history into the model dimension via a hidden-to-latent projector, and predicts the next planning state. The Decoder then aggregates these states using an Exponential Moving Average (EMA) mechanism to create stabilized aggregators, which are finally projected via a decoder projector to synthesize the coarse-grained reasoning step.
Training and Refinement
The model is trained during Supervised Fine-Tuning (SFT) via a reconstruction loss calculated as the crossentropy between the ground-truth text and the model’s prediction conditioned on the state Sk. To improve robustness, Gaussian Noise is injected into accumulated states during training.
For policy refinement, Reinforcement Learning (RL) is employed to optimize only the Decoder parameters while freezing all Planner parameters to maintain structural integrity of the learned latent manifold. A Decoupled GRPO objective is used, where diversity is introduced by enabling temperature sampling in the Decoder from fixed latent states Sk. The reward function provides dense supervision based on semantic validity and mathematical correctness for both intermediate steps and the final answer.
Key Findings and Trade-offs
Empirical results reveal a distinct trade-off between greedy precision and exploration potential. While PLaT achieves lower greedy accuracy than baselines, it demonstrates superior scalability in terms of reasoning diversity, exhibiting a steeper scaling slope in Pass@k metrics. This indicates that PLaT learns a robust, broader solution space rather than overfitting to a narrow trajectory.
Analysis of latent states confirms the interpretability advantage: PLaT allows for analysis of the reasoning topology by comparing branching characteristics against explicit CoT. The model maintains significantly higher entropy throughout the majority of the reasoning process compared to explicit CoT and CODI, suggesting that latent states do not collapse to a single mode but rather maintain a superposition of multiple potential verbalizations until the final termination signal is required.
Efficiency and Hyperparameter Analysis
PLaT achieves significant speedup compared to Explicit CoT, reducing inference latency by approximately 56% by skipping intermediate token generation. The efficiency results show that PLaT provides a favorable balance, delivering transparent, high-diversity reasoning at a speed significantly faster than standard CoT.
Sensitivity analysis indicates that the optimal configuration involves NL = 2 for latent sequence length and αEMA = 0.5 for the EMA coefficient when NL = 2. A dimension of ds = 2048 serves as a robust sweet spot across all benchmarks,
suggesting that while increasing dimensions does not yield substantial gains, it increases computational overhead. The framework enables dynamic termination via Lazy Decoding, where reasoning is paused if the first token generated by the Decoder is not the answer delimiter (tans
).
Conclusion and Limitations
PLaT reformulates latent reasoning as planning over latent space, shifting from implicit pattern matching to planning in continuous space. Its primary contribution is enabling dynamic termination of latent planning and intermediate interpretability of latent states.
A limitation noted is that RL refinement may cause performance degradation on OOD tasks, suggesting the policy overfits to the training domain's reward signal. Future work should investigate joint optimization strategies for the Planner and Decoder, explore scaling laws beyond NL=2, and validate its efficacy in less-structured domains like creative writing.
C Prompts for LLM Judgments
The paper also details prompts designed to leverage LLMs for reasoning validation, including a prompt for clustering semantically equivalent steps and a prompt for validating the mathematical correctness of individual reasoning steps. These prompts are crucial for the RL reward function, where an automated evaluator (GPT-4o-mini) is used to verify step validity with high consistency.
D Limitations and Future Work
While PLaT introduces a promising paradigm for decoupled latent planning, there are several limitations in our current implementation that outline directions for future research.
Improvements for AI systems
Here are the specific improvements that can be made to current AI systems by implementing the PLaT framework, along with what the resulting improved system can achieve:
Improvement 1: Transition from Rigid Chain-of-Thought (CoT) to Dynamic Latent Search Planning.
The core improvement is replacing the discrete token path collapse of standard CoT with a continuous latent manifold planning process. Instead of forcing the model to commit to a single, potentially suboptimal token at every step, PLaT models reasoning as an autoregressive trajectory in high-dimensional continuous space (the Planner).
The improved AI system can:
-
Perform
System 2
style deliberation by maintaining a superposition of multiple potential reasoning strategies simultaneously within the latent space. -
Dynamically decide when to terminate reasoning based on semantic probes (Lazy Decoding), rather than relying on fixed hyperparameters, leading to more efficient inference for complex tasks.
Improvement 2: Implement Decoupled Planning and Verbalization Architecture.
The separation of the Planner (reasoning) and the Decoder (verbalization) allows for specialized optimization of each component. The Planner evolves deterministic latent states, while the Decoder focuses solely on grounding these thoughts into natural language via a reconstruction objective.
The improved AI system can:
-
Achieve superior scalability in reasoning diversity (Pass@k metrics) because it learns a broader solution space rather than overfitting to a narrow path.
-
Enable interpretable intermediate reasoning by decoding the latent states back into text on demand, allowing for post-hoc analysis of the model's
thought process.
Improvement 3: Optimize Policy Refinement via Decoupled Reinforcement Learning (GRPO).
By freezing the Planner and only training the Decoder parameters using a Group Relative Policy Optimization (GRPO) objective, we decouple planning stability from exploration. The RL signal is used to refine the linguistic policy (the mouth
) without distorting the underlying reasoning logic.
The improved AI system can:
-
Achieve better greedy accuracy on in-domain tasks by training the Decoder to maximize reward signals derived from correct formatting and mathematical correctness, while preserving a stable reasoning topology.
-
Be more robust against reward hacking compared to end-to-end RL methods because the core reasoning mechanism (the Planner) is structurally preserved.
Improvement 4: Leverage Latent State Analysis for Reasoning Topology Understanding.
The ability to decode latent states allows researchers and developers to analyze the internal structure of the model's reasoning decisions. This involves using external tools like GPT-4o-mini for clustering semantically equivalent reasoning steps and validating their mathematical validity.
The improved AI system can:
- Provide a
glass-box
mechanism where the complex, continuous latent states are mapped back to discrete, verifiable logical steps, allowing for rigorous auditing of its decision-making process against ground truth.
Improvement 5: Adaptive Hyperparameter Selection for Reasoning Depth.
By analyzing the sensitivity of the latent sequence length (NL) and the EMA coefficient (αEMA), researchers can determine optimal configurations tailored to specific benchmarks (e.g., NL=2 for GSM8k).
The improved AI system can:
- Be deployed with a configuration that balances precision and exploration optimally for a given problem type, avoiding the performance degradation seen when excessively long reasoning chains are used.
This PLaT-based system will be significantly more capable of tackling complex, multi-step mathematical or logical problems than current state-of-the-art CoT models by trading a small amount of deterministic greedy accuracy for vastly superior exploration and scalability in finding diverse, correct solutions.
Sources
- GPT-4 Technical Report
- Evaluating Large Language Models Trained on Code
- Reasoning Beyond Language: A Comprehensive Survey on Latent Chain-of-Thought Reasoning
- Training Verifiers to Solve Math Word Problems
- Implicit Chain of Thought Reasoning via Knowledge Distillation
- Efficient Reasoning Models: A Survey
- PAL: Program-aided Language Models
- GPT-4o System Card
- From System 1 to System 2: A Survey of Reasoning Large Language Models
- Large Language Model Guided Tree-of-Thought
- ProphetNet: Predicting Future N-gram for Sequence-to-Sequence Pre-training
- Your LLM Knows the Future: Uncovering Its Multi-Token Prediction Potential
- CODI: Compressing Chain-of-Thought into Continuous Space via Self-Distillation
- Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models
- System-1.5 Reasoning: Traversal in Language and Latent Spaces with Dynamic Shortcuts
- SoftCoT: Soft Chain-of-Thought for Efficient Reasoning with LLMs
- Beyond Chain-of-Thought, Effective Graph-of-Thought Reasoning in Language Models
- Soft Thinking: Unlocking the Reasoning Potential of LLMs in Continuous Concept Space
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection