Better World Models Can Lead to Better Post-Training Performance
cs.LG, cs.AI
Submitted: 2025-12-03
Updated: 2026-08-29
Code: https://github.com/prakharg55/CubeLM-NeurIPS-MI
License: http://creativecommons.org/licenses/by-sa/4.0/
The gist: We study how explicit world-modeling objectives affect the internal representations and downstream capability of Transformers, using Rubik's Cubes as our training domain.
Terminology
Abstract
We study how explicit world-modeling objectives affect the internal representations and downstream capability of Transformers, using Rubik's Cubes as our training domain. We ask: (1) how does explicitly pretraining a world model affect a model's latent representations, (2) how does world-model quality affect post-training performance, and (3) how should a finite data budget be split between pretraining and task-specific fine-tuning? We compare standard action-prediction fine-tuning with two strategies that add explicit state-prediction supervision: state pretraining followed by fine-tuning, and joint action and state training. We measure task accuracy after Group Relative Policy Optimization (GRPO). We further find that explicit world-modeling yields better representations in terms of higher probing accuracy and steerability of the model, and that better representations yield larger gains from GRPO, especially on harder cube states. Finally, when the fine-tuning budget is held fixed, probe accuracy strongly predicts the GRPO improvement. Under a fixed total data budget, accuracy is maximized by allocating only a small fraction to pretraining.
Sources
- Understanding intermediate layers using linear classifier probes
- Language Models Represent Space and Time
- World Models
- A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity
- Shared Global and Local Geometry of Language Model Embeddings
- What Does it Mean for a Neural Network to Learn a "World Model"?
- Emergent Linear Representations in World Models of Self-Supervised Sequence Models
- The Geometry of Categorical and Hierarchical Concepts in Large Language Models
- The Linear Representation Hypothesis and the Geometry of Large Language Models
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks