The Objective Is the Bottleneck: Latent World Models Encode What Their Planners Cannot Use

arXiv:2608.12959 · cs.LG, cs.AI · Submitted 2026-08-13 · Read on arXiv

Joyjeet Singh

cs.LG, cs.AI

Submitted: 2026-08-13

Updated: 2026-08-14

Comments: Follow-up to arXiv:2608.10145. All experiments run on a laptop CPU; no model was trained or fine-tuned. Code, checkpoints and every measurement: github.com/joyjeet-singh/tinylab

Code: https://github.com/joyjeet-singh/tinylab

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: This paper investigates why latent world models fail at long-horizon planning, using a reproduction of LeWorldModel on the TwoRoom environment.

Terminology

Summary

This paper investigates why latent world models fail at long-horizon planning, using a reproduction of LeWorldModel on the TwoRoom environment. The central finding is that the bottleneck is the planner's objective function, not the predictor's quality.

The predictor is not the bottleneck. Rolling the released checkpoint forward autoregressively on real validation clips shows that at fifteen planner steps (75 environment steps), the imagined state is still only 0.189 as wrong as assuming the world froze, with error growing slowly and smoothly. The planner only ever imagines 25 environment steps (horizon 5 with frameskip 5), so the model can see three times farther than it is ever asked to.

The objective saturates and inverts. The planner uses cross-entropy-method (CEM) planning that minimizes squared latent distance between the imagined final embedding and the goal embedding. Measuring pairwise latent distance against true distance across 7,140 pairs of real recorded positions reveals: overall Pearson r = 0.426; within-band correlation is 0.668 under twenty units but below 0.13 beyond forty units; the metric saturates across 80–300 units with mean varying by only −3.3%; and past about 120 units the relationship turns negative, meaning moving away from the goal can lower the planner's cost. This explains observed failures: 26 of 37 offset-100 failures finish farther from the goal than they began at a median 1.41× the original separation, and success falls off exactly where the metric goes blind (100% under 60 units, 37% between 60–100, 6% between 100–140).

The pathology belongs to the method, not one reimplementation. The same measurement on the original authors' released checkpoint shows Pearson 0.388, Spearman 0.423, non-monotone, with −2.2% spread across 80–300 units. Across all four checkpoints, long-horizon planning success rank-orders exactly with metric quality and inversely with one-step prediction accuracy. The most accurate predictor is the worst long-horizon planner because accuracy and metric usability move in opposite directions in these runs. The split tracks training pipeline differences (history size 1 and action dim 2 for usable geometry versus history size 3 and dense ten-wide actions for pathological ones), but these factors are confounded and cannot be separated.

The information is present. A ridge probe fit on frozen embeddings of 350 rendered positions and evaluated on 150 held-out positions recovers position at R2 = 0.9922, with decoded-position distance monotone at r = 0.9897 and +73.4% spread across 80–300 units, versus latent L2 at r = 0.437, non-monotone, +1.2% spread. The same holds on the authors' checkpoint (probe R2 = 0.9627). The encoder represents what the planner needs; the objective discards it.

The repair. Replacing only the objective, with nothing retrained and no GPU, lifts goals reached at offset 100 from 26.0% to 98.0% using a learned temporal-distance cost, equals the 98.0% reached at offset 25, and still reaches 92.0% under a third of the budget (50-step budget instead of 150). Planning stops depending on the horizon. A decoded-position cost reaches 88.0% at offset 100.

Reachability, not proximity, is the right objective. A cost learned from temporal separation alone (how many steps apart two observed frames were, supervised by frame separation on disjoint episode splits) predicts spatial distance worse than a position probe (r = 0.819 against 0.9897) yet plans better (98.0% against 88.0%). Testing on matched spatial distance pairs split by whether they lie in the same room or across the dividing wall: the temporal head charges 24% more to cross the wall (cross/same ratio 1.24), while squared latent distance charges 4% less (ratio 0.96). The decoded-position cost cannot represent the wall at all (ratio 1.00 by construction). Planning success across the three objectives orders exactly by cross/same cost ratio (0.96 → 1.00 → 1.24) rather than accuracy against spatial distance (0.437 → 0.9897 → 0.819).

Where the repair fails. On the authors' released weights, the learned cost is beaten by the linear decoded-position cost: 34.0% against 70.0%. The reason is that a head fit on encoded real frames is evaluated on imagined ones, and the authors' predictor drifts far enough off the encoding manifold for an MLP to extrapolate (degradation +74% versus +3% on the authors' versus our checkpoint). Training the head on both real and imagined pairs (rolling the predictor forward under recorded block-mean actions) removes the mismatch (+11% rather than +74%) and lifts planning to 34.0%, but does not close the gap. The prescription is conditional: a learned reachability cost wins where the predictor is accurate enough (one-step error 0.116), while a linear cost is the robust choice where it is not (one-step error 0.410).

Limitations. One environment (TwoRoom), one seed per checkpoint, confounded training conditions, no claim about the cause of saturation, and the repair is a planner change that leaves the model's own objective untouched.

Improvements for AI systems

Improvements to AI systems:

  1. Replace latent-distance planning objectives with learned temporal-reachability costs. Instead of minimizing squared L2 distance in latent space (which saturates and inverts beyond 40 units), train a small head to predict temporal separation between observed frames, and use that as the planner's cost. This yields 98% goal-reaching vs 26% on offset-100 tasks, and removes horizon-dependence.

  2. Add a linear decoded-position cost as a robust fallback. When the world model's predictor drifts off the encoding manifold (detectable via one-step prediction error > 0.3), switch to a linear probe on decoded positions. This achieves 70% success where the learned cost fails (34%), and is more robust to distribution shift.

  3. Train the cost head on both real and imagined rollouts. To prevent evaluation mismatch, generate imagined states by rolling the predictor forward under recorded block-mean actions, and include those pairs in the head's training data. This reduces degradation from +74% to +11% and improves planning by 3×.

  4. Use reachability-aware costs that encode environmental topology. The learned temporal cost naturally charges more for crossing walls (cross/same ratio 1.24) while latent distance charges less (0.96). This makes planning respect obstacles and connectivity, not just Euclidean proximity.

  5. Add a metric-quality diagnostic before deployment. Before using a planner, measure the correlation between the cost function and true task-relevant distance (e.g., Pearson r, monotonicity, spread across relevant ranges). If r < 0.5 or the relationship is non-monotone, replace the objective. This prevents silent failures in long-horizon tasks.

What the improved AI system can do:

  • Plan reliably over long horizons (e.g., 100+ environment steps) without performance collapse, reaching 98% success vs 26% previously.

  • Remain robust to world-model drift by automatically selecting between learned reachability and linear decoded-position costs based on predictor accuracy.

  • Generalize across checkpoints and training pipelines without retraining the world model or encoder, using only a planner-side change.

  • Respect environmental structure (walls, obstacles, connectivity) in planning, enabling navigation that avoids impossible paths.

  • Provide early warning of objective failure via a cheap diagnostic, allowing systems to switch strategies before task failure.

Abstract

Latent world models are judged by how well they predict, so when planning fails at long horizons the natural reading is that the predictor degrades. On a reproduction of LeWorldModel on TwoRoom we show the binding constraint is the planner's objective instead. The predictor is not the limit: its imagined state seventy-five environment steps ahead is still only 0.189 as wrong as assuming the world froze, while the planner never imagines beyond twenty-five. The objective is. Cross-entropy-method planning minimises squared latent distance, which tracks true distance at r = 0.426, saturates by about eighty arena units and decreases beyond a hundred and twenty, so moving away from the goal can lower the cost. The information is present throughout: a ridge probe recovers position from the frozen embedding at R squared 0.9922. The pathology is the method's, not one reimplementation's. It is present in the authors' released weights, and across four checkpoints long-horizon success rank-orders exactly with metric quality and inversely with prediction accuracy. Replacing only the objective, with nothing retrained and no GPU, lifts goals reached at offset 100 from 26.0% to 98.0%, equals the 98.0% at offset 25, and reaches 92.0% under a third of the budget: planning stops depending on the horizon. The best cost is not the most accurate. A head learned from frame separation alone predicts spatial distance worse than a position probe (r = 0.819 against 0.9897) yet plans better, charging 24% more to cross the environment's dividing wall where squared latent distance charges 4% less. It has learned reachability, not proximity.

Sources

Related papers