A Unifying Perspective on Causal World Models: From Observations to Representations to Structure

arXiv:2608.13456 · cs.AI, cs.CV · Submitted 2026-08-13 · Read on arXiv

Imperial College London

cs.AI, cs.CV

Submitted: 2026-08-13

Updated: 2026-09-03

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: This paper studies world models (WMs) from a causal perspective across multiple levels of abstraction, ranging from perceptual observations to building a conceptual representation of the structure

Terminology

Summary

This paper studies world models (WMs) from a causal perspective across multiple levels of abstraction, ranging from perceptual observations to building a conceptual representation of the structure governing environment dynamics. The authors argue that useful WMs must go beyond generative capabilities alone: they should also capture entity properties, entity-to-entity interactions, and entity-to-environment interactions that determine and explain the dynamics of a system. The paper provides a formal definition of Causal WMs (CWMs) grounded in the tasks they are intended to support, connecting world modelling with existing work in causal representation learning, object-centric learning, causal discovery, structural causal models, and model-based decision-making. Finally, the paper relates CWMs to the literature on identifiability, clarifying when the components of a WM can be recovered from data and up to which equivalence.

The paper begins by noting that World models have become one of the central aspirations of modern Artificial Intelligence and that the challenge is to build such models from observations, trajectories, interventions, and domain knowledge, while avoiding two unhelpful extremes: treating WMs as an unstructured black box with no explicit variables or mechanisms, or assuming in advance that the right variables and mechanisms have already been specified. The paper notes that recent work occupies different points along this spectrum, including predictive world models such as Dreamer, hierarchical variants such as ResDreamer, joint-embedding approaches such as LeJEPA, and causal dynamics methods.

The paper's central problem is "how to turn observations into a CWM by combining three strands of work: causal representation learning, addressing when latent variables can be recovered from unstructured observations; causal discovery, studying when cause-effect relationships over the structured variables (represented by Gr) can be recovered; and model-based decision-making, covering when actions and utilities can be used for controlled behaviour. The paper presents a conceptual overview of a Causal Ladder of CWM" as a four-level progression from perception to intervention and imagination, following Pearl's causal hierarchy.

The paper makes two contributions. First, it formalises a CWM as a Markov decision process that links observations, latent states, actions, transition distributions, and utility. Second, it argues that guarantees for CWMs should be understood component-wise, each imposing different assumptions and admitting different admissible equivalences, so that the right notion of identifiability for CWMs is the one that preserves the downstream reasoning task and the corresponding notion of control.

The paper formalises the notion of relational variables in Definition 1, where for each ordered pair of entities i, j in an entity set O, the relational variable components are defined, with diagonal blocks retaining entity attributes and off-diagonal blocks encoding interactions. Definition 2 defines the formal state description as st = (xt, rt) where xt is the observation and rt is the structured relational state. Definition 3 defines a CWM as a tuple W = (X, A, RO O, P, U), where X is the observation space, A is the action space, RO is the relational latent-state space for entity set O, and U is a utility function over actions and states. The action-conditioned observation-level transition factorises as P(xt+1 = x′ xt = x, at = a) = ∫ P(xt+1 rt+1)P(rt+1 rt, at = a)P(rt xt) drt, rt+1.

The paper's Remark 1 separates a CWM into components that must be learned or specified: an inference model from observations to relational variables, a transition or intervention model over relational variables, an action model, a prediction model back to observations, and a utility function. The paper states that A WM can support causal reasoning only if the variables that jointly drive prediction and utility are available to the model, either as recorded components of the observation or as latent factors inferred from it, making the standard causal sufficiency assumption.

For identifiability, the paper defines strong identifiability (Definition 4) and identifiability up to equivalence (Definition 5). The paper notes that Strong identifiability is rarely the right target for CWMs because entity permutations and invertible latent reparameterisations can leave the learnt mechanisms unchanged. Table 1 makes the equivalence relation concrete across the CWM's representation, distributions, structure and utility, distinguishing representation equivalences for the latent state, entity blocks, entity-attribute blocks, and inter-entity relational blocks from compatibility conditions on the CWM interfaces.

The paper concludes that CWMs should be understood as structured decision models rather than as monolithic predictors and that the central contribution is a component-wise view, where each part of the model may be identifiable only up to admissible equivalence that enables the downstream reasoning. This decomposition clarifies what can be learned from data, what must be supplied by interventions or domain knowledge, and when predictive models can support causal reasoning. Future work should turn this perspective into a rigorous formalisation, followed by practical algorithms and empirical testing, while extending it to partial observability and latent confounding.

Improvements for AI systems

Improvements to AI Systems:

  1. Component-wise Identifiability-Aware World Models
  • Build world models that explicitly track which components (inference, transition, action, prediction, utility) are identifiable from data and up to which equivalence class (e.g., permutation, reparameterization).

  • The AI system can automatically flag when its latent state or relational structure is only known up to an admissible transformation, preventing overconfident causal claims and guiding when to request interventions or domain knowledge.

  1. Relational Causal State Representations
  • Implement latent states as structured relational variables (entity attributes + pairwise interactions) rather than flat vectors, following Definition 1.

  • The AI system can reason about entity-to-entity and entity-to-environment interactions explicitly, enabling compositional generalization to new object counts, unseen entity pairings, and counterfactual what-if queries about specific interactions.

  1. Causal Ladder of World Models for Decision-Making
  • Train world models that progress through four levels: perception → structured latent state → intervention-aware transitions → imagination/planning under counterfactuals (Pearl’s ladder).

  • The AI system can distinguish between observational predictions, action-conditioned predictions, and counterfactual reasoning, allowing it to answer what would happen if I had done X instead of Y in model-based reinforcement learning, not just what happens next.

  1. Utility-Aware Causal Abstraction
  • Integrate the utility function directly into the world model definition (as in Definition 3), so that the learned latent variables are optimized not only for prediction accuracy but also for downstream decision value.

  • The AI system can learn compressed causal states that preserve only the information relevant for control, reducing sample complexity and improving robustness to irrelevant perceptual details.

  1. Equivalence-Class-Conditioned Planning
  • Use the identifiability results (Table 1) to define planning algorithms that operate over equivalence classes of latent states rather than point estimates.

  • The AI system can plan robustly even when the latent representation is only known up to a transformation, by making decisions invariant to those admissible reparameterizations—leading to more reliable policies in partially observed or non-identifiable settings.

  1. Causal Sufficiency Checking and Active Intervention
  • Implement a module that checks whether the current latent variables satisfy causal sufficiency (i.e., all common causes are captured). If not, the system can propose targeted interventions to disambiguate confounders.

  • The AI system can actively query its environment (or a simulator) to reduce identifiability gaps, improving its causal model’s reliability for planning and explanation.

  1. Hierarchical Causal World Models for Long-Horizon Tasks
  • Extend the CWM framework to hierarchical variants (e.g., ResDreamer-style) where higher levels capture slow-changing causal mechanisms and lower levels capture fast dynamics.

  • The AI system can plan at multiple temporal abstractions, using causal structure at each level to avoid compounding errors and to transfer knowledge across tasks with shared causal mechanisms.

  1. Object-Centric Counterfactual Imagination
  • Combine object-centric learning with the CWM’s relational state to enable counterfactual simulations that swap, remove, or modify specific entities and their interactions.

  • The AI system can generate novel hypothetical scenarios (e.g., what if this object were heavier?) for robust policy training, data augmentation, and explainable AI—without needing real-world trials.

  1. Identifiability-Guided Data Collection
  • Use the formal identifiability conditions to design data acquisition strategies (e.g., which interventions to perform, which observations to collect) that minimize the equivalence class of the learned CWM.

  • The AI system can efficiently allocate its exploration budget to resolve the most uncertain causal mechanisms, accelerating learning in sparse-reward or high-stakes environments.

  1. Causal World Model as a Structured Decision Model
  • Replace monolithic predictive models with modular CWMs where each component (inference, transition, action, utility) can be independently updated, validated, or replaced using domain knowledge.

  • The AI system can incorporate expert-provided causal mechanisms for known parts (e.g., physics) while learning unknown parts from data, leading to safer and more sample-efficient deployment in real-world robotics or scientific discovery.

Abstract

World Models (WM) are increasingly seen as a foundation for intelligent agents that can predict, plan, and act beyond their training distribution. In this paper, we study WMs from a causal perspective across multiple levels of abstraction, ranging from perceptual observations to building a conceptual representation of the structure governing the environment dynamics. We argue that useful WMs must go beyond generative capabilities alone: they should also capture entity properties, entity-to-entity interactions, and entity-to-environment interactions that determine and explain the dynamics of a system. We provide a formal definition of Causal WMs (CWMs) grounded in the tasks they are intended to support, connecting world modelling with existing work in causal representation learning, object-centric learning, causal discovery, structural causal models, and model-based decision-making. Finally, we relate CWMs to the literature on identifiability, clarifying when the components of a WM can be recovered from data and up to which equivalence. With this, we ground WMs in representations and structures that support causal reasoning and informed decision-making.

Sources

Related papers