Textual Planning with Explicit Latent Transitions
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Textual Planning with Explicit Latent Transitions".
Jane: Planning with large language models is bottlenecked by token-by-token generation and repeated full forward passes, making multi-step lookahead and rollout-based search expensive in latency and compute.
Tom: First, who's behind it and why it matters.
Paper summary: Tom: Well, we've been looking at the paper "Textual Planning with Explicit Latent Transitions," and it seems like they are tackling a real issue in how we use large language models for planning. The main idea is that token-by-token generation and running full forward passes repeatedly really slow down multi-step lookahead and rollout based search, which makes things expensive in terms of both speed and compute resources.
Jane: It sounds like the paper introduces EMBEDPLAN as a way to bypass that bottleneck by using a lightweight transition model instead of generating the next state step-by-step. So, essentially, they are trying to make planning much faster by moving away from autoregressive generation.
Lu: What's really fascinating is how they frame the problem by encoding natural language state and action descriptions into vectors and then using a transition network to predict the next-state embedding. It suggests we can operate within a frozen language embedding space, which is quite a novel approach.
Meng: From an engineering standpoint, I wonder how effective this frozen embedding approach is in capturing the necessary dynamics for planning without needing extensive fine-tuning on every new problem. It sounds promising for deployment if it keeps the model size manageable.
Lalam: I'm seeing how powerful this concept is; if we can learn these latent transitions, it could dramatically improve how our AI handles complex, multi-step reasoning tasks by making the search process much more efficient.
Tom: Exactly! So, the core claim of "Textual Planning with Explicit Latent Transitions" is that they can achieve fast planning computation without having to finetune the encoder for every new problem. They propose replacing autoregressive next-state generation with a lightweight transition model that predicts the next-state embedding, and this is what makes it potentially so much faster.
Jane: It really boils down to using contrastive learning objectives—specifically state prediction and action disambiguation—to train this network. This training teaches the model to correctly identify next states among candidates and to distinguish the effects of different actions on the same state.
Lu: That contrastive learning setup is key; it’s not just about predicting a single outcome, but also understanding fine-grained action semantics through that action disambiguation loss. This suggests a richer understanding of the planning problem structure itself.
Meng: I'm curious about the data they used, though they mention training on a cumulative dataset of over three million transitions across nine classical PDDL planning domains from ACPBENCH. Getting that kind of breadth in the data is impressive for establishing those initial dynamics.
Lalam: Having that massive dataset means they’ve built up a really robust understanding of how different states evolve across various planning scenarios, which is crucial for any AI system.
Paper summary: Tom: Right, and the results show that when they test this on interpolation within known problem manifolds, they get near-perfect next-state retrieval at ninety-nine point seven percent Hit@five. That's a very strong starting point for seeing how well the system learns the patterns.
Jane: That near-perfect interpolation performance suggests that within problems they’ve seen before, the method is highly effective at retrieving correct next states. It shows the latent transition learning is working well when it's operating within established problem structures.
Lu: But the real interest lies in how they characterize this capability boundary; they found a sharp capability boundary for latent transition learning in frozen embedding spaces. This implies there's a limit to what these frozen embeddings can support beyond the manifolds they were trained on.
Meng: So, that sharp boundary is where things get tricky; it means we can do great work inside the known boundaries, but moving outside those boundaries leads to a significant drop in performance. That's something I need to consider for practical application.
Lalam: That sharp boundary is a critical piece of information because it tells us that the current limitation isn't just a lack of data, but really about the representation itself. It points toward a fundamental issue with how we're encoding knowledge in these frozen spaces.
Tom: And that leads us perfectly into the next part of the paper, which discusses how they test this generalization using six different protocols. They test everything from interpolation and plan variants to extrapolation and cross-domain transfer.
Jane: Testing those different splits is how they map out where the system actually succeeds and where it starts to struggle. It’s a comprehensive way to see the limits of the EMBEDPLAN approach.
Lu: And when we look at the cross-domain transfer results, they show performance as low as six point six percent Hit@five for zero-shot domain transfer. That's a very low number when you compare it to the baseline.
Meng: A six point six percent Hit@five suggests that trying to take this learned knowledge and apply it to an entirely different set of problems without any overlap is not feasible right now. That means for real-world, diverse applications, we can't just throw a single model at everything.
Lalam: That low cross-domain performance reinforces what they suggest about the limitation; it confirms that the embeddings are encoding lexical surface form instead of the structural role in planning. It's not learning transferable planning logic yet.
Tom: That distinction is really important; they are saying that acquiring domain-specific transition knowledge, rather than generalizing within domains, is the main sticking point right now. But they did show that action disambiguation helps a lot with in-domain generalization.
Paper summary: Jane: So, while the system struggles to jump between different problem domains, improving the action disambiguation loss actually boosts its performance within a single domain by about three point four times for Action Acc@five. That's a tangible win for focused planning tasks.
Lu: The way they combine the state prediction loss and the action disambiguation loss—setting λ to two after grid search—is what seems to drive that performance boost by nineteen point three percentage points in Hit@five. That tuning process is quite methodical.
Meng: From a practical standpoint, I mean if we use that combined loss, we should see better performance on tasks that require precise action selection within a known planning structure. It suggests focusing on the quality of action semantics is where the immediate gains are.
Lalam: If we can improve action disambiguation, it could lead to AI systems that are much more reliable when making sequential decisions in complex scenarios. That reliability is what we need for practical use.
Tom: So, to wrap up the summary of "Textual Planning with Explicit Latent Transitions," the paper shows that EMBEDPLAN can achieve near-perfect interpolation within known problem manifolds but struggles significantly when forced to generalize to unseen problems or different domains.
Jane: It really highlights that the current limitation in using frozen embeddings for this type of learning is that they capture surface form rather than the underlying structural role in planning. This suggests we need a way to teach the model what actions *do* structurally, not just what words describe them.
Lu: The big picture here is that this work establishes a very clear capability boundary for this type of latent transition learning in frozen embedding spaces. It tells us exactly where the current approach stops providing useful predictive power without retraining the encoder.
Meng: If we're building systems that need to adapt quickly to completely new operational environments, this limitation is a major hurdle for now. We need something that handles structural shifts better.
Lalam: And that's what makes this paper so interesting; it doesn't just show a result, it clearly demarcates the current frontier of what frozen embedding space planning can do versus where we need to go next. It directs future research very specifically.
Tom: And that's what we've got for this segment on "Textual Planning with Explicit Latent Transitions," focusing on how it moves beyond simple next-state generation by using latent transitions and contrastive learning. Next up, we’ll talk about what these findings mean for the future of AI planning in our next segment.
Conclusion: Tom: So we’ve been diving deep into how these new methods are trying to make AI planning much more efficient by using latent transitions, and now we’re wrapping up our discussion on this paper called "Textual Planning with Explicit Latent Transitions."
Jane: Exactly, Tom. It sounds like the core idea is that instead of generating every single step one by one, the AI uses a smarter model to predict the overall change between states, which cuts down on those slow, repetitive calculations.
Lu: I think what’s really striking is how they managed to put this whole concept into a framework where you can use frozen embeddings from models like Llama or Qwen without needing massive retraining for every new planning problem. That capability alone opens up so many creative doors for exploring complex search spaces.
Meng: From an engineering standpoint, I’m focused on the practicality of that frozen space; if the model relies only on what it learned in training, how robust is it when faced with a completely novel planning task?
Lalam: I see huge potential here because if we can get this kind of structural understanding into our models, we could create systems that learn to reason about how things change in complex cultural or operational contexts far more naturally.
Tom: That's the big picture, Lalam. And looking at the authors, it seems they really focused on building a solid methodology around contrastive learning objectives to get this transition prediction right.
Jane: It’s interesting that they used those specific loss functions—state prediction and action disambiguation—to guide the learning process; that shows a very deliberate approach to teaching the AI what matters most in planning.
Lu: That intentional design is crucial because it hints at a deeper understanding of how we should represent planning knowledge, moving beyond just surface-level text matching.
Meng: I’m still thinking about scaling this up; if the framework works well for a few domains, how does that translate when we have hundreds of them? We need to know if this is something that can truly handle the complexity of real-world scenarios without losing fidelity.
Lalam: The implication for our culture is that we might see AI systems capable of more nuanced, long-term strategic planning, which could fundamentally alter how we approach complex decision-making processes across many fields.
Tom: That’s exactly what this paper suggests—we're moving toward models that can navigate complexity much more smoothly than before.
Jane: It really moves the conversation away from just generating text and toward building systems that understand the underlying logic of state transitions.
Lu: I think this work sets a very clear benchmark for what latent transition learning can achieve when constrained by frozen embeddings, which is super valuable context for all future research in this area.
Meng: So, while it shows great internal consistency within known problem manifolds, the paper clearly flags that the transfer outside those bounds is still quite limited right now.
Lalam: That limitation actually makes the next step even more important because it tells us precisely what kind of structural knowledge we need to acquire before we can hope for true cross-domain understanding.
Technion – Israel Institute of Technology
cs.CL
Submitted: 2026-02-04
Updated: 2026-10-01
Importance score: 83/100
The gist: Planning with large language models is bottlenecked by token-by-token generation and repeated full forward passes, making multi-step lookahead and rollout-based search expensive in latency and
Key concepts
- EMBEDPLAN Framework
- This framework replaces traditional autoregressive generation with a lightweight transition model. It uses frozen language embeddings and a learned network to predict the next state embedding based on text descriptions of the current state and action, allowing for faster transitions.
- Frozen Embeddings Space
- Instead of training large models from scratch, this method uses pre-trained language encoders (like Llama or Qwen) whose weights are frozen. These embeddings map natural language descriptions into a shared 128-d space where the transition model learns to operate efficiently.
- Action Disambiguation Loss (Laction)
- This specific training objective teaches the model to distinguish between different actions that yield the same current state. By optimizing this loss, the model learns fine-grained semantic differences in planning actions, which significantly boosts performance within a single problem domain.
Terminology
Summary
Planning with large language models is bottlenecked by token-by-token generation and repeated full forward passes, making multi-step lookahead and rollout-based search expensive in latency and compute. The gist: frozen embeddings can support transition learning within known problem manifolds but fail to generalize beyond them, establishing a sharp capability boundary where near-perfect interpolation occurs, meaningful extrapolation is achievable within a domain, but zero-shot cross-domain transfer remains nearly absent.
EMBEDPLAN Framework
The proposed framework introduces EMBEDPLAN, which replaces autoregressive next-state generation with a lightweight transition model operating in a frozen language embedding space. This involves encoding natural language state and action descriptions into vectors, training a lightweight network to predict the next-state embedding, and retrieving the next state by nearest-neighbor similarity. The training combines two contrastive objectives: state prediction,
which learns to identify correct next states among candidates, and action disambiguation,
which learns to distinguish the effects of different actions applied to the same state.
Architecture and Components
The EMBEDPLAN architecture comprises three main parts: (1) a frozen encoder that maps text descriptions into embeddings, (2) learned projection heads that reduce dimensionality to a shared 128-d space, and (3) a transition network predicting the next-state embedding. The model utilizes four different frozen encoders: MPNet, BGE-M3, Qwen2.5-7B, and Llama-3.3-70B. The transition network is trained using a composite contrastive objective:
-
The state prediction loss (Lstate) is InfoNCE over next-state candidates to capture coarse domain dynamics.
-
The action disambiguation loss (Laction) distinguishes the effects of different actions applied to the same state, teaching fine-grained action semantics.
Data and Training Objectives
The model is trained on a cumulative dataset of over 3 million transitions across 9 classical PDDL planning domains sourced from ACPBENCH. For each problem instance, trajectories are extracted to form state-action-next-state triplets. The training objective is defined as:
(Equation 3)
L = Lstate + λ · Laction
where λ controls the emphasis on action disambiguation, set to 2 after grid search. The training utilizes a batch size of 128, a temperature τ = 0.07, and AdamW optimization with a learning rate of 4 × 10−5.
Generalization Protocols
The study rigorously characterizes generalization boundaries using six protocols:
-
Interpolation split: Measures interpolation within known problem manifolds where transitions are randomly partitioned between train and test sets. This protocol shows
near-perfect next-state retrieval (99.7% Hit@5 under Interpolation).
-
Plan-Variant split: Evaluates full-plan execution on alternative optimal plans for problems seen during training, testing generalization to
unseen solution paths under fixed problem structure.
-
Extrapolation split: Assigns entire problems to train or test sets, measuring generalization to
new problem configurations within a known domain.
-
Cross-Domain transfer: Evaluates zero-shot domain transfer by training on one source domain and testing on a different target domain with no shared problems or structure, showing performance as low as
6.6% Hit@5.
-
Multi-Domain learning: Trains a single model on all 9 domains simultaneously to test
capacity sharing without catastrophic forgetting.
-
Leave-One-Out (LOO) generalization: Tests transfer from diverse training by training on 8 domains and evaluating on the held-out 9th domain, achieving only
9.2% Hit@5.
Key Findings on Generalization
The results reveal a sharp capability boundary for latent transition learning in frozen embedding spaces.
Within known problem manifolds, EMBEDPLAN achieves high performance. However, generalization to unseen problems degrades substantially: performance drops to 54.6% on unseen problem configurations (Extrapolation; a 45.2 pp gap).
Furthermore, transfer across domain boundaries fails, with Cross-Domain reaching only 6.6% Hit@5 (+2.7 pp above the 3.9% untrained baseline).
The study attributes this failure to embeddings encoding lexical surface form rather than the structural role in planning.
While action disambiguation improves within-domain generalization, it learns domain-specific semantics that do not transfer across domain boundaries. The gap structure confirms that acquiring domain-specific transition knowledge, not generalizing within domains, is the primary bottleneck.
Action Disambiguation Impact
The action disambiguation loss significantly boosts performance. Training with the full composite objective (λ = 2) improves Action Acc@5 by 3.4× and Hit@5 by 19.3 pp compared to training with state prediction loss only (λ = 0).
Improvements for AI systems
Here are specific improvements to AI systems based on the EMBEDPLAN framework:
-
Replacement of token-by-token autoregressive generation for planning with a lightweight, frozen transition model operating in a fixed language embedding space. This fundamentally reduces latency and compute required for multi-step lookahead and rollout-based search, making long-horizon planning computationally feasible at inference time.
-
Implementation of an explicit latent transition function learned via contrastive objectives (state prediction and action disambiguation). The system can predict the next state embedding from the current state embedding and action embedding, enabling fast nearest-neighbor retrieval for next states without needing to re-invoke a large language model for every step.
-
Deployment of a decoupled architecture where semantic understanding is handled by frozen LLM encoders (e.g., Llama or MPNet embeddings), while dynamics prediction is handled by a small, learned transition network (<500K parameters). This allows for modular updates of the dynamics model independently from the expensive language encoder.
-
Enabling
Plan-Variant
generalization: The system can be trained to recognize and execute alternative optimal action sequences for a given problem instance, rather than merely memorizing a single trajectory. This allows the AI to adapt its plan based on slight variations in action ordering or choice while maintaining the same underlying problem structure. -
Support for within-domain dynamics learning: Once exposed to a specific planning domain (e.g., Blocksworld), the system can learn and utilize its transition model to efficiently solve new, unseen problems within that same domain by extrapolating learned dynamics from observed trajectories.
-
Improved action semantics in planning: By incorporating an action disambiguation loss, the system learns to precisely distinguish the causal effect of different actions applied to a state (e.g., knowing that
pick-up(A)
is distinct fromstack(A,B)
). This leads to more robust and less error-prone planning sequences. -
Quantification of generalization limits: The system can be rigorously characterized by evaluating its performance across six protocols (Interpolation, Plan-Variant, Extrapolation, Cross-Domain Transfer, Multi-Domain Learning, and Leave-One-Out). This provides a clear diagnostic boundary for when the frozen embedding space is sufficient versus when domain transfer or structural generalization is required.
This improved AI system can perform:
-
Fast inference for complex planning tasks in known domains (e.g., logistics, block manipulation) with significantly reduced latency compared to current LLM-based planners.
-
More reliable plan generation that can handle slight variations in optimal action sequences within the same problem structure.
-
Efficient learning of new problems within a familiar domain after initial exposure, without needing retraining of the entire language model encoder.
-
Robust decision-making by accurately distinguishing between actions with similar surface-level descriptions but different underlying causal effects.
Sources
- Relational inductive biases, deep learning, and graph networks
- LLM+P: Empowering Large Language Models with Optimal Planning Proficiency
- Representation Learning with Contrastive Predictive Coding
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering