SAGE: Subgoal-Conditioned Action Generation for Latent World Model Planning
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "SAGE: Subgoal-Conditioned Action Generation for Latent World Model Planning".
Tom: Latent world models are powerful planning paradigms that have struggled with proposal quality as planning horizons grow, and this paper introduces SAGE,
Jane: First, who's behind it and why it matters.
Paper summary: Tom: So, we’re talking about SAGE, which is essentially a prior-conditioned planner that uses latent subgoal decomposition to structure how it searches for actions. Jane That means instead of just throwing random action proposals at the world model, the planner gets guidance based on predicted intermediate goals.
Lu: What this paper claims is that they decompose long-horizon tasks into reachable intermediate states using a variable-length latent subgoal generator, which conditions the generation of candidate action sequences for a frozen world model. This captures both fine-grained local dynamics and higher-level task progress through subgoals of varying durations.
Meng: From an engineering standpoint, replacing random initialization with structured guidance sounds like it directly addresses the issue where proposal quality struggles as planning horizons grow. How do they actually achieve this structural conditioning?
Tom: Well, SAGE uses two learned components for this. First, there’s a latent subgoal generator that predicts the expected latent state at a future time based on the current history and the far goal tokens. Then, they have a subgoal-conditioned action generator that maps that predicted target to a distribution of action options.
Jane: That’s really clever because the subgoal generator is parameterized by a four-layer Transformer decoder designed to produce local targets that the frozen world model can evaluate against. It’s not just guessing; it’s predicting what the next meaningful step in terms of latent progress should look like.
Lalam: I see this as a cultural improvement because if we can structure our planning based on these predicted subgoals, it means the AI isn't just reacting locally; it understands the overall task progression better, which could lead to more reliable and predictable behavior in complex environments.
Lu: The paper mentions that shorter subgoals keep local control fine while longer ones capture higher-level progress toward the final goal, allowing one generator to support multiple subgoal lengths. This flexibility is what makes the framework powerful for reasoning over different temporal scales.
Meng: I'm interested in the training part; how do they train that subgoal generator to produce those meaningful targets? Does it just rely on expert trajectories, or is there a specific loss function?
Tom: They align windows from expert trajectories where a window contains history, a far goal at t plus delta, a local future at t plus tau, and the tau expert actions between them. The subgoal generator is then trained to match the frozen world model's latent representation of that local future.
Jane: And for the action generator, they use a trajectory mixture negative log-likelihood loss, supervised by the expert trajectory itself. So, you’re training both parts to work together based on real expert demonstrations.
Lalam: It suggests that we can improve the culture of AI development by making the planning process more structured and goal-oriented rather than purely reactive, which could lead to more intentional system design.
Conclusion: Tom: So, we’re wrapping up our chat on SAGE, focusing on the core idea that this method couples latent subgoal decomposition with prior-conditioned action generation to significantly improve long-horizon planning while keeping strong short-horizon performance.
Jane: The authors, Latian Cheng and Qi Zhang from Peking University, basically show how using variable-length subgoals as priors directly shapes the candidate distribution explored by the planner. It’s about using those predicted targets to guide the search process instead of starting from a generic random distribution.
Lu: The implication here is that we can extend the planning capability of existing latent world models without needing to modify or retrain their predictive dynamics, just by adding this structural layer. It shows that improving the structure of action proposals has a direct impact on planning ability even with frozen models.
Meng: I’m thinking about the practical impact; if this method helps us handle much longer planning horizons effectively, it means AI systems could tackle more complex, multi-stage tasks in real-world scenarios where sequential decision-making is key.
Lalam: For our culture, this points toward a future where the AI’s planning isn't just about immediate next steps but about maintaining a coherent vision of the entire task progression, which is a much more robust way to build reliable systems.
Tom: It really boils down to this: by generating local targets and using them as proposal priors, we see substantial gains in planning success when the target offset is large—for instance, pushing PushT success from twelve point seven percent up to sixty-four point seven percent when the target offset hits one hundred fifty.
Jane: And it’s important to remember that while this method is powerful for long horizons, the paper does acknowledge that LeWM-based refinement is still essential for selecting and refining the final action sequence after SAGE generates those structured proposals.
Lu: That distinction between generating a proposal structure and having a world model evaluate it remains key, showing that the two parts are complementary rather than one replacing the other.
Meng: So, the main implication is that we can leverage existing latent world models more effectively by adding this structured guidance mechanism without having to completely overhaul their underlying predictive dynamics.
Lalam: It suggests a path toward building AI agents whose planning is inherently hierarchical and goal-aware, which really improves the overall reliability and structure of those systems we develop.
Tom: That’s a lot to process, but the core message from SAGE is that structuring how we propose actions based on predicted subgoals unlocks much better performance across various task complexities.
Peking University
cs.AI
Submitted: 2026-07-20
Updated: 2026-09-28
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 85/100
The gist: Latent world models are powerful planning paradigms that have struggled with proposal quality as planning horizons grow, and this paper introduces SAGE, a prior-conditioned planner that uses latent
Key concepts
- Latent World Model (LeWM)
- A frozen, low-dimensional representation of the world learned from expert data. It acts as a simulator that can evaluate imagined future trajectories without needing to re-train its dynamics, allowing the planner to test potential actions against realistic outcomes.
- Latent Subgoal Generator
- A four-layer Transformer decoder that predicts a specific latent state ($\hat{z}_{t+\tau}$) representing an expected intermediate goal at a future time. It takes current history and goal information as input to guide the search toward meaningful, reachable milestones.
- Subgoal-Conditioned Action Generator
- This component maps the current situation and the predicted latent subgoal to a distribution of possible actions. By conditioning action generation on these subgoals, it ensures that proposed action sequences are explicitly linked to achieving specific intermediate targets.
- CEM Refinement
- A search refinement technique used during planning. After generating initial action proposals, CEM selects the most promising region based on evaluation by the frozen LeWM and refines the best proposal through repeated sampling rounds to select a high-quality final action sequence.
Terminology
Summary
Latent world models are powerful planning paradigms that have struggled with proposal quality as planning horizons grow, and this paper introduces SAGE, a prior-conditioned planner that uses latent subgoal decomposition to structure action search. The gist: SAGE decomposes long-horizon tasks into reachable intermediate states using a variable-length latent subgoal generator to condition the generation of candidate action sequences for a frozen world model.
How it works
The core idea is to replace random proposal initialization with structured guidance by conditioning actions on predicted latent subgoals rather than initializing from a generic random distribution. This is achieved through two main learned components: a variable-length latent subgoal generator and a subgoal-conditioned action generator, both operating on a frozen world model.
- The latent subgoal generator takes the current observation history, the low-dimensional state, the far goal tokens, embeddings of the remaining goal offset, and a requested duration τ as input to predict the expected latent state at time t + τ:
zˆt+τ = zg t+δ + Dψ(zt−k:t, xt, zg t+δ, δ, τ). This generator is parameterized by a four-layer Transformer decoder designed to produce a duration-matched local latent target that the frozen world model can use for evaluation.
- The subgoal-conditioned action generator maps the current history, far goal, and generated target to a distribution of action options. This component takes the history tokens, low-dimensional state, predicted local subgoal (zˆt+τ), far-goal tokens, and temporal variables as input to produce a trajectory-level Gaussian mixture qϕ(at:t+τ−1 zt−k:t, xt, zˆt+τ, zg t+δ, δ, τ). This mixture allows the planner to generate options of different lengths while maintaining an explicit correspondence between an option and its predicted subgoal.
Training and Planning Process
The training process involves aligning windows from expert trajectories where a window contains a history, a far goal at t + δ, a local future at t + τ, and the τ expert actions between them. The subgoal generator is trained to match the frozen LeWM latent of the local future: Lsubgoal = SmoothL1(ˆzt+τ, zt+τ) + λcos [1 − cos(ˆzt+τ, zt+τ)]. The action generator is trained using a trajectory mixture negative log-likelihood loss, where the desired action sequence is supervised by the expert trajectory.
During online planning, a query with total horizon H is decomposed into a duration schedule (τ1,..., τR). At each stage r, SAGE predicts zˆt+τr from the current history and remaining goal. It then samples K action options of duration τr and gives them to the frozen LeWM. The frozen LeWM evaluates these imagined futures against the generated local target (zˆt+τr), and CEM refines the best proposal region before execution, following a shared 30 refinement round procedure.
Experimental Validation
Experiments on PushT and OGBench Cube validate that coupling latent subgoal decomposition with prior-conditioned action generation substantially improves long-horizon planning while preserving strong short-horizon performance. When the target offset H=150, this method raises PushT success from 12.7% to 64.7% and OGBench Cube success from 26.7% to 67.3%. The results show that the benefit accumulates as a route contains more replanning stages, and that the world-model-guided search remains essential for selecting and refining the final action sequence, as evidenced by SAGE prior top reaching 16.0% at H=150 compared to 64.7% for the complete planner.
Key Contributions
The contributions of this work can be summarized as:
** We introduce a multi-horizon latent subgoal generator that decomposes distant goals into reachable intermediate states at different temporal horizons, capturing both fine-grained local dynamics and higher-level task progress.**
(We develop a subgoal-conditioned action generator that generates action candidates conditioned on each predicted subgoal and uses a frozen latent world model to evaluate and refine these proposals.)
(Experiments on PushT and OGBench Cube show that combining variable-length subgoal generation with subgoal-conditioned action generation substantially improves long-horizon planning while preserving strong short-horizon performance.)
Furthermore, the study shows that both generating local targets and using them as proposal priors yield substantial gains, while LeWM-based refinement remains essential for selecting the final action sequence. The temporal schedule study also highlights that both the duration and ordering of local decisions affect planning outcomes. The results suggest that improving the structure of action proposals can extend the planning capability of existing latent world models without modifying or retraining their predictive dynamics.
Improvements for AI systems
Based on this scientific paper, here are the specific improvements that can be made to AI systems by implementing the SAGE (Subgoal-Conditioned Action Generation) framework:
-
Enhanced Long-Horizon Planning Capability: The primary improvement is a substantial increase in planning success rates for long-horizon tasks (e.g., PushT and OGBench Cube).
-
Improved Goal Alignment Across Temporal Scales: The system can effectively decompose distant goals into a sequence of temporally relevant intermediate subgoals, balancing fine-grained local control with high-level task progress.
-
Structured Action Proposal Generation: Instead of relying on generic random action initialization for the proposal distribution, the AI will sample actions conditioned on predicted latent subgoals. This concentrates the search space onto plausible behaviors relevant to achieving the immediate local objective while maintaining a direction toward the distant goal.
-
Robustness to Long Planning Horizons: The system maintains strong short-horizon performance (e.g., H=25) while showing significant gains as the target offset increases (up to H=150), overcoming the
proposal quality
constraint that plagues standard latent world model planners at large horizons. -
Improved Sequential Decision Making via Temporal Abstraction: The ability to reason over multiple temporal scales allows the AI to coordinate decisions across different time horizons within a single planning stage, leading to more coherent and efficient long-horizon trajectories.
-
Adaptability through Variable-Duration Scheduling: The planner can utilize different temporal schedules (e.g., repeated short commitments vs. longer, mixed sequences) during online execution, allowing the agent to adapt its replanning frequency based on the current planning stage's needs (as shown in Table 3).
-
Preservation of Learned Dynamics: A key benefit is that these improvements are achieved without modifying or retraining the underlying frozen latent world model (LeWM); only the components responsible for target prediction and action conditioning are learned and optimized.
By implementing SAGE, an AI system can perform complex robotic manipulation or navigation tasks that require sustained, multi-step planning where traditional planners fail due to poor exploration of the exponentially large action space when aiming for distant objectives.
Abstract
Latent world models have emerged as a powerful planning paradigm by learning action-conditioned predictive dynamics and using them as internal simulators to imagine and evaluate candidate action sequences. However, as the planning horizon grows, performance becomes increasingly constrained by proposal quality: a fixed candidate budget must search an exponentially larger action space, making it difficult to expose the world model to high-quality candidate futures for evaluation. In this paper, we introduce SAGE, a prior-conditioned planner that replaces random proposal initialization with structured guidance. At each planning stage, a goal-conditioned generator predicts the next intermediate latent subgoal for a specified duration, which is then used to condition the generation of candidate action sequences. To capture semantic information across temporal scales, we use subgoals of varying durations as priors, balancing fine-grained local control with higher-level long-horizon progress. Then the frozen world model evaluates these proposals against the same subgoal and guides their refinement before execution. Experiments on PushT and OGBench Cube show that coupling latent subgoal decomposition with prior-conditioned action generation substantially improves long-horizon planning while preserving strong short-horizon performance. To be specific, when the target offset is 150, it raises PushT success from 4.7% to 64.7% and OGBench Cube success from 20.7% to 67.3%. We further extend latent world-model planning to LIBERO, where SAGE improves full-episode success from 0% with the vanilla LeWM planner to 48.7% on Scene2 and 58% on Caddy.
Sources
- V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
- Beyond the Next Step: Variable-Length Latent World Models for Long-Horizon Planning
- Mastering Diverse Domains through World Models
- LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels
- FF-JEPA: Long-Horizon Planning in World Models with Latent Planners
- PRISM: PRior-guided Imagination Sampling in world Models
- Hierarchical Planning with Latent World Models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection