How do World Models and Policies Compose in LLM Agents? A Joint Spectral and Behavioral Account
cs.LG, cs.AI, cs.CL
Submitted: 2026-08-30
Updated: 2026-08-30
Comments: Accepted to EMNLP 2026
Code: https://github.com/Rick-Xu315/WM_Spectral
License: http://creativecommons.org/licenses/by/4.0/
The gist: How do LLM agents come to both understand environments they act in and master tasks set within them? Through controlled experiments combining world-model training (next-state prediction) and policy
Terminology
Abstract
How do LLM agents come to both understand environments they act in and master tasks set within them? Through controlled experiments combining world-model training (next-state prediction) and policy training (reward maximization), we investigate this question. We dissect the resulting models through their additive parameter updates. Geometrically, we find effective world-model updates are low-rank and share an input-feature subspace with policy updates while writing to nearly orthogonal output directions, whether trained separately or sequentially. However, we find that, in projection interventions, the sequential update induces more robustness than separate policy RL when removing the world model's leading input directions, suggesting that it has learned alternative input pathways. Behaviorally, we find the sequentially trained agent explores a wider range of states and actions. Based on this, we ask: does policy training preserve world knowledge as well as it could? We probe this with training-free merging built on the geometrically motivated input basis plus an online world-model loss during policy RL, and show both improve over the untreated baseline. Our findings suggest world knowledge and task-directed ability can be learned in geometrically complementary forms, and that future post-training pipelines should consider how best to engineer the interface between them.
Sources
- Mastering Diverse Domains through World Models
- Evaluating Large Language Models Trained on Code
- Why Do LLM Agents Fail in Exploring New Environments? A World-Modeling Perspective
- Locating and Editing Factual Associations in GPT
- Qwen2.5 Technical Report
- Qwen3 Technical Report
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Dyna-Mind: Learning to Simulate from Experience for Better AI Agents
- Dyna-Think: Synergizing Reasoning, Acting, and World Model Simulation in AI Agents
- ECHO: Terminal Agents Learn World Models for Free
- Reinforcement World Model Learning for LLM-based Agents
- Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models
- Representation Engineering: A Top-Down Approach to AI Transparency
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks