The Evaluation Protocol Determines the Result: An Independent Reproduction of LeWorldModel on TwoRoom
Joyjeet Singh
cs.LG
Submitted: 2026-08-10
Updated: 2026-08-12
Comments: Independent reproduction of arXiv:2603.19312 - https://github.com/joyjeet-singh/tinylab
Code: https://github.com/joyjeet-singh/tinylab
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 75/100
The gist: The paper reports an independent reproduction of LeWorldModel on the TwoRoom diagnostic environment.
Terminology
Summary
The paper reports an independent reproduction of LeWorldModel on the TwoRoom diagnostic environment. The authors reimplemented the method from released code and paper, trained four models on rented GPUs at roughly six dollars each, and ran all evaluation on a single laptop CPU.
The representation claim reproduces directly: a ridge probe fitted on 80% recovers position at R2 0.9977 (Pearson r 0.9988), and a two-layer network on the same split reaches 0.9994
against a reported Pearson correlation of approximately 0.996.
The planning claim reproduces under one of two published evaluation protocols. The released material contains two conflicting protocols: Appendix F.1 specifies a 150-step budget with the goal sampled 100 frames ahead; the released evaluation configuration specifies 50 and 25.
On the authors' own released weights, these yield 14.0% and 84.0% respectively, and only the configuration's values reproduce the reported figure. The authors' checkpoint reaches 42/50 = 84.0%
under the repository's evaluation configuration, against a reported 87%. The authors' corrected checkpoint reaches 47/50 = 94.0%
at the repository's goal offset of 25 steps, with a 95% Wilson interval of [83.8%, 97.9%] that contains the reported figure. However, changing nothing but how the goal is constructed moves that same checkpoint from 84.0% to 8.0%
on fifty identical episodes (matched-pair p < 10−8).
The training claim does not reproduce from released configuration files alone. Four undocumented conventions were required: "actions must be gathered densely across a frameskip block rather than sub-sampled, the action-encoder width is set programmatically from that block, pixels are ImageNet-normalised, and actions are z-scored by dataset statistics with NaN rows removed. A reproducer following released configurations alone
obtains a model whose predictor cannot converge" — its training loss plateaus at 0.30, while the corrected pipeline descends monotonically to a held-out value 36 times lower (0.0086 vs 0.3064).
Two findings generalize beyond the reproduction. First, one-step prediction accuracy does not predict long-horizon planning success
: across three checkpoints spanning a sevenfold range in prediction error, accuracy orders short-horizon success monotonically (78.0%, 84.0%, 94.0% at offset 25) but fails to order long-horizon success at all (54.0%, 12.0%, 20.0% at offset 100 under a 50-step budget), where the two most accurate checkpoints finish farther from the goal than a random-action policy does
(mean final distances of 122.5 and 116.6 units vs 111.1 for random). Second, a batch normalisation layer inflated our reported validation loss by a factor of up to 300, concealing for three training runs a training loss that was flat or descending throughout
— the artifact requires both a large amplification (running variance far below activation scale) and weights still moving.
The paper also reports a pre-registered mechanism-level result that did not survive: a same-room planning advantage of +39.1 points at p = 3.4 × 10−8 on one checkpoint fell to +12.7 points under a change of action scaling and to −6.4 points on a different checkpoint. The authors conclude that effect sizes measured on a single reproduction checkpoint should not be read as properties of the method.
Additional findings include: a 32-pixel debugging fixture that looked visually near-identical to the real environment sat twenty-five times farther from the training distribution than training frames sit from each other
and produced below-random planning results while a position probe read R2 0.99; the effective rank of the embedding cloud rose from 18.6 to 67.8 of 192 dimensions under the corrected pipeline while carrying the same position and action information; and the latent geometry does represent the wall (cross-wall pairs separated by 1.79 times same-room pairs at matched physical distance).
The paper's limitations include a single seed throughout, ten training epochs rather than the repository's one hundred, and a non-empty deviation set documented in Table 1. All code, checkpoints, evaluation reports, fidelity audit, pre-registration, and gate outputs are released at github.com/joyjeet-singh/tinylab.
Improvements for AI systems
Improvements to AI systems:
-
Add protocol-consistency validation to evaluation harnesses. Before reporting any planning metric, automatically detect and flag conflicting evaluation configurations (e.g., step budgets, goal offsets) across code, configs, and appendices. The improved system refuses to emit a single number unless all protocols agree or explicitly report both.
-
Implement sensitivity auditing for goal construction. For any model that plans toward a goal, run a matched-pair test across at least two goal-encoding variants (e.g., one-hot vs. coordinate, scaled vs. unscaled). The system reports a
goal-robustness score
and flags results that swing more than 10% under trivial encoding changes, as seen here (84% → 8%). -
Add a
prediction-vs-planning decoupling
diagnostic. For any world model, train checkpoints at multiple training epochs and compute both one-step prediction error and long-horizon planning success. The system automatically warns if the two metrics are not monotonically aligned, preventing false confidence from low prediction loss. -
Integrate batch-normalization artifact detection. During training, monitor the ratio of running variance to activation scale and the magnitude of weight updates. If a BN layer's running stats diverge from activations by >100× while weights are still moving, the system flags the validation loss as untrustworthy and recomputes it with BN in eval mode or after freezing.
-
Build a distribution-distance probe for debugging fixtures. For any synthetic or simplified environment used in testing, compute the MMD or Wasserstein distance between its frame distribution and the training distribution. The system rejects fixtures whose distance exceeds, say, 10× the typical inter-training-frame distance, preventing misleading
near-identical
visual checks. -
Require multi-checkpoint effect-size reporting. Any claim of a mechanism-level advantage (e.g., same-room planning) must be averaged over at least three independently trained checkpoints with different random seeds and action-scaling variants. The system automatically computes a confidence interval and flags if the effect changes sign across checkpoints.
-
Add a
reproducibility gate
for released configs. When a model is released, the system runs a minimal training from those configs alone and checks convergence within a factor of 2 of the reported loss. If it fails (as here, plateauing at 0.30 vs. 0.0086), it blocks the release and demands documentation of all undocumented conventions.
What the improved AI system can do:
It can produce planning metrics that survive trivial changes in goal encoding, distinguish genuine long-horizon ability from short-horizon prediction accuracy, avoid being misled by BN-induced loss inflation, reject visually similar but distributionally distant test environments, and report effect sizes that generalize across seeds and scaling choices—all while automatically auditing its own evaluation protocols for consistency.
Abstract
LeWorldModel trains a latent world model with a prediction loss and a single anti-collapse regulariser, and reports approximately 87% of goals reached on TwoRoom, its simplest diagnostic environment. We reproduce that result by independent reimplementation on roughly 25 of rented compute, with all evaluation on one laptop CPU. We reach 94.0% at the repository's evaluation goal offset, against 84.0% for the authors' own released checkpoint measured under our protocol on identical episodes, and we reproduce the reported representation result directly (position probe Pearson r = 0.9988 against a reported 0.996). Reaching that point required four conventions that determine the outcome and appear in no released configuration file: dense action gathering across a frameskip block, a programmatically-set action-encoder width, ImageNet pixel normalisation, and action z-scoring. A reproducer following the released configurations alone obtains a model whose predictor cannot converge. The evaluation protocol is itself contested by the released material. The paper's appendix and the repository's configuration specify different goal offsets and step budgets; on the authors' own weights these yield 14.0% and 84.0%, and only the configuration's values reproduce the reported figure. On fifty identical episodes, changing nothing but how the goal is constructed moves that checkpoint from 84.0% to 8.0%. Two findings generalise. One-step prediction accuracy does not predict long-horizon planning success: across three checkpoints spanning a sevenfold range in prediction error, including the authors' own, it orders short-horizon success monotonically and fails to order long-horizon success at all. And a batch normalisation layer inflated our reported validation loss by up to a factor of 300, concealing a training loss that was flat throughout.
Sources
- LeJEPA: Provable and Scalable Self-Supervised Learning Without the Heuristics
- World Model Control by Trajectory Reachability Metrics
- stable-worldmodel-v1: Reproducible World Modeling Research and Evaluation
- LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks