Can Video World Models Track Unobserved World States?
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "Can Video World Models Track Unobserved World States?".
Tom: Video world models are increasingly used as simulators, yet visual fidelity alone does not show that a model maintains the hidden state of the world.
Jane: First, who's behind it and why it matters.
Paper summary: Tom: Hey team, I’m really pumped about this paper we’re talking about today because it tackles a fundamental issue with video world models: whether they can actually remember what's hidden from view.
Jane: It sounds like the core idea is that just having good visuals isn't enough; the model needs to maintain a consistent internal understanding of the scene even when things are occluded or unseen, right?
Lu: Exactly. This paper, "Can Video World Models Track Unobserved World States?", looks at this using an action-conditioned video Shell Game as a way to test how these models handle tracking tasks like S5 state tracking.
Meng: So, if we look at the abstract from this paper, what's the main problem they're setting out to solve with this study?
Tom: Well, the paper points out that current video model backbones trained only on visible content struggle to carry unobserved states when new actions or observations arrive later on. The authors examine this gap by using an action-conditioned video Shell Game as a visual version of S5 state tracking.
Jane: So, essentially they're asking if these models can remember the ball's position even when it’s hidden across several swaps in that game?
Lu: That’s the main thesis: they show that standard architectures—like autoregressive Transformers or Mamba-based models—can fit a training horizon of five swaps, but they fail to generalize robustly when the swap chain gets longer.
Meng: So, what's the specific reason why they fail on longer chains? Is it just a limitation of the architecture itself?
Tom: Not exactly; they find that this failure happens because the pixel-based diffusion target doesn't supervise that unseen hidden state; instead, those generated frames can't carry it and the state has to live inside the architecture rather than in the tokens.
Jane: That means when you look at a Transformer, its architectural state is just an append-only KV cache, which leads to each chunk having to re-derive the arrangement from the whole history at a fixed depth.
Lu: Right, and they identify two specific mechanisms that allow models to extrapolate beyond their training length. First, they look at widening the recurrent transition for linear attention variants from zero one to
−one one: , which allows for reflection-like transitions with negative eigenvalues necessary for composing a swap.
Paper summary: Meng: And the second mechanism involves a Test-Time Training fast weight whose online nonlinear feature-map update makes its effective kernel history-dependent, moving it outside plain linear attention by rewriting the SwiGLU inner model.
Tom: So, while these mechanisms help them fit the training horizon of five swaps, they still fall toward chance on longer swap chains and provide no benefit when rendering plausible video with extra denoising steps.
Jane: It seems like even with these architectural tweaks, the pixel-based diffusion target never supervises the unseen hidden state in a way that helps the model learn to carry it forward.
Lu: The study also looked at probing internal states by saving model readouts and using fixed-bank readouts as queries into the KV cache or input to those fast weights, and they found that while visible swap identity decodes at one point zero, hidden-ball accuracy remains quite weak across the models they tested.
Meng: That suggests that we might be using these readouts mostly for rendering the current view rather than actually carrying a compact persistent state needed for tracking.
Tom: It really emphasizes that the current setup means every serial step is spent on rendering, and we observe no scratchpad emerging when you look at the rollout in DiT, unlike how chain-of-thought in language models works with working memory.
Jane: So the paper suggests that for a video model to be useful beyond short sequences, it needs a state carried across chunks paired with an update rule expressive enough to compose that hidden transition on that state in place.
Lu: That points toward needing architectures where the update isn't just additive, but history-dependent, and where the model can correct its state from observations as well.
Meng: From an engineering standpoint, if we have to re-derive the arrangement from history at every chunk for Transformers, that sounds computationally expensive when you scale up to complex scenarios.
Tom: It is expensive if you don't have a compact way to store and revise that state during the sequence generation process.
Jane: So, what does this imply for how we design video world models moving forward? Does it suggest a shift away from purely additive updates toward something more integrated?
Paper summary: Lu: The paper suggests that reliable tracking in these settings will likely require progress in stateful architectures and objectives that reach beyond just the visible content.
Meng: I wonder if this means we need to prioritize developing mechanisms like those fast weights or recurrent transitions over simply increasing the number of tokens for memory.
Tom: Precisely, because the findings show that simply adding scratchpad tokens doesn't guarantee success up to longer swap chains; LaCT performed better than plain additive updates because it normalized states via those fast weights and kept extrapolating as a better-regularized register.
Jane: So, the implication is that the focus needs to shift from just fitting the training length to creating architectures that inherently support state composition across unseen parts of a sequence.
Lu: We are seeing evidence that current AR video backbones can't compose unobserved state updates, which means they should fail whenever a correct frame depends on a long swap chain among five or more hidden objects.
Meng: If we want to deploy these in complex simulators, the challenge isn't just rendering plausible frames; it's ensuring that internal logic remains consistent over long interactions.
Tom: So, this paper gives us clear direction: we need state carried across chunks paired with an update rule expressive enough to compose the hidden transition on that state in place.
Jane: It’s a call for architectures where the update isn't merely additive but history-dependent, and where the model can correct its state from observations as well.
Lu: These findings have broad implications for dynamic world exploration tasks like Memory Maze and three dee Block World, which require a state that is not just a pure function of the action stream.
Meng: I think we need to look at how we integrate observation updates directly into the state transition mechanism, rather than treating them as just another input to the rendering pipeline.
Tom: It really highlights that while LaCT is one of the strongest backbones across most settings, achieving that correction from observations remains an open challenge for video world models.
Jane: So, we’re looking at a future where reliable tracking will likely require progress in stateful architectures and objectives that reach beyond just visible content.
Conclusion: Segment: Conclusion — The Big Picture** **(Recap)** So, to wrap up our discussion on this paper, we've seen how they tested video world models with a shell game and found that current methods struggle when the hidden state gets too complicated or long.
Tom: I think the title itself really captures the essence of the problem they are solving here—can these models actually keep track of what's happening behind closed doors in a dynamic environment? The authors did a solid job showing exactly where those models break down when it comes to memory.
Jane: It’s fascinating because it moves past just making pretty videos and gets into the core challenge of intelligence: maintaining an internal model of reality, even when you can't see everything. They show that the visual information alone isn't enough to sustain that internal knowledge over time.
Lu: From a theoretical standpoint, the authors are pushing us toward a new way of thinking about video processing, suggesting we need recurrent transitions or explicit memory structures because simple autoregressive flows just aren't structured to handle this kind of sequential state dependency effectively.
Meng: I look at the practical implications here and see that if we can't reliably track those unobserved states, then any complex simulation involving long interactions or hidden objects in a physical space will be fundamentally limited by the model’s inability to maintain a consistent world state.
Lalam: If this research moves us toward architectures that can genuinely compose hidden transitions in place, it opens up avenues for much richer cultural simulations where AI agents can learn and remember complex social or physical scenarios more deeply.
Tom: Exactly, and I think the authors' findings suggest that we need to look beyond just making the next frame look good; we need to build systems where the model is actively managing a persistent internal representation of the world.
Jane: It really boils down to this: for video models to be truly useful in dynamic exploration, they have to become stateful entities that can update their knowledge based on what they observe, not just blindly generate what looks right next.
Lu: That means the future direction involves integrating observation updates directly into the state transition mechanism so the model can correct its internal representation when it gets new input.
Meng: That correction ability is crucial; if a model drifts from its internal world state, it loses track of reality in a complex scenario, which is a major hurdle for deployment in anything beyond simple demonstrations.
Lalam: I think this work paves the way for AI systems that possess more robust "working memory," allowing them to handle long-term planning and context far better than we see now. **(Hook)** This brings us perfectly to how these concepts translate into actual design changes, which is what we'll explore next with our engineering lead.
Joonghyuk Shin, Yicong Hong, Jaesik Park, Xun Huang
Seoul National University · Roblox
cs.CV
Submitted: 2026-08-31
Updated: 2026-09-28
Comments: Project webpage:https://joonghyuk.com/stateful-vwm-web/
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
Importance score: 77/100
The gist: Video world models are increasingly used as simulators, yet visual fidelity alone does not show that a model maintains the hidden state of the world.
Key concepts
- State Tracking
- This is the ability of a world model to remember what is not currently visible in the video and correctly update that memory as new actions and observations occur. It ensures the model can predict future frames based on past hidden information, not just what is currently seen.
- Recurrent Transition
- This refers to how a model's internal state evolves from one time step to the next. The paper suggests this needs to be expressive enough—like a nonlinear RNN—to allow the model to correctly compose complex transitions for hidden objects, rather than just adding new information linearly.
- Test-Time Training (TTT) Fast Weights
- This is a mechanism where the model's update process changes during inference. By making this update depend on its own previous state, it creates an effective kernel history. This allows the model to behave differently at test time than it did during training, helping it maintain a persistent internal state.
- Autoregressive KV Cache
- This is a memory structure used in Transformer models to store past information. The paper argues this cache is not a compact running state; instead, it forces the model to re-derive the arrangement from the entire history every time, limiting its ability to carry a simple, persistent hidden state.
Terminology
Summary
Video world models are increasingly used as simulators, yet visual fidelity alone does not show that a model maintains the hidden state of the world. The gist: current video model backbones supervised solely on visible content cannot carry unobserved states, which makes recent discussions on state tracking directly relevant: how the state is implicitly carried across chunks, and how expressive its transition is.
The Problem of State Tracking in Video World Models
A useful world model must remember what is no longer visible, update it as new actions and observations arrive, and recover it when it becomes visually relevant again. When the current frame is only a partial observation (e.g., an occluded object, a camera turned away, or a hidden arrangement transformed by actions), the next correct frame is determined not by the present image but by a state carried and updated across the intervening steps. Recent theoretical work suggests that robust state tracking requires reintroducing serial computation, either through a sufficiently expressive recurrent transition (like a nonlinear RNN) or through chain-of-thought (CoT) scratchpad tokens.
Empirical Study on the Shell Game
The paper examines an action-conditioned video Shell Game as a visual analogue of S5 state tracking. The Shell Game renders the swap-only S5 task of compositional state tracking as action-conditioned video, where the ball remains unobserved across a sequence of swaps. The core finding is that standard backbones—autoregressive Transformers, SSM (Mamba)-based models, and bidirectional Transformers—fit the training horizon but fail to extrapolate robustly when the swap chain grows beyond a fixed depth. This failure occurs because the pixel-based diffusion target never supervises the unseen hidden state; instead, the generated frames cannot carry it and the state has to live inside the architecture rather than in the tokens.
Mechanisms Enabling State Extrapolation
The study identifies two mechanisms that allow models to extrapolate beyond their training length:
-
A recurrent transition widened to admit negative eigenvalues. For linear attention variants, widening the recurrence from [0, 1] to [−1, 1] allows for
reflection-like transitions
with negative eigenvalues, which are necessary for realizing a swap composition. -
A Test-Time Training (TTT) fast weight whose online nonlinear feature-map update makes its effective kernel history-dependent, placing it outside plain linear attention. This is achieved by rewriting the SwiGLU inner model such that the update reaches the feature map through which it reads its own state, making the
effective kernel history-dependent.
Architectural and Ablation Insights
The analysis of different architectures and mechanisms reveals specific dependencies:
-Autoregressive KV Cache:
-Autoregressive KV cache is a visual history, not a compact running state. For autoregressive Transformers, the KV cache cannot update state in place; it can only write the new state at a higher layer where it becomes readable. This leads to each chunk has to re-derive the arrangement from the whole history at a fixed depth.
-Probing internal states:
The paper probes internal states by saving model readouts and using fixed-bank readouts (a bank of random Gaussian vectors) as an attention query into the KV cache or as input to TTT fast weights. These probes show that while the current readout decodes visible swap identity at 1.0, hidden-ball accuracy remains far weaker
across models, suggesting the readout is used primarily for rendering rather than carrying a compact persistent state.
-State Carriers:
The explicit scratchpad token layer (adding M memory tokens per layer) lifts the baseline to perfect in-distribution accuracy and generalizes up to N=10 swaps, but accuracy drops to chance by N=20 swaps.
LaCT performs better than the plain additive update because it better normalizes states via fast weights and keeps extrapolating as a better-regularized register.
Implications for General World Models
The Shell Game experiments suggest that current AR video backbones cannot compose unobserved state updates, implying they should fail whenever a correct frame depends on a long swap chain among five or more hidden objects. The paper concludes that reliable tracking requires a state carried across chunks paired with an update rule expressive enough to compose the hidden transition on that state in place.
This highlights the need for architectures where the update is not merely additive but history-dependent, and where the model can correct its state from observations as well.
Future Directions
The work discusses broader implications for dynamic world exploration tasks like Memory Maze and 3D Block World. These harder cases require a state that is not a pure function of the action stream: it must also be updated from observations.
The paper suggests that while LaCT is one of the strongest backbones across most settings, the need to correct state from observations remains an open challenge for video world models. Reliable tracking will likely require progress in stateful architectures, and objectives that reach beyond visible content.
Improvements for AI systems
Based on the provided research paper, here are specific, actionable improvements for AI systems and what those improved systems could achieve:
The core improvement is shifting video world models from being plausible
generators to being stateful trackers
by integrating mechanisms that allow them to maintain a hidden configuration across unobserved frames.
Here are the specific improvements derived from the paper:
-
Enhance Video World Models with Expressive State Transition Mechanisms:
-
Implement Test-Time Training (TTT) with Nonlinear Fast Weights for Online State Revision:
-
Integrate Explicit Token-Space Registers (Scratchpads) for Persistent State Carrying:
-
Utilize Negative Eigenvalue Linear Attention Variants for Compositional Tracking:
The resulting improved AI systems could perform the following specific tasks:
-
Improved long-horizon tracking of hidden objects in complex environments (e.g., tracking the ball in the Shell Game) across sequences significantly longer than their training horizon (extrapolation).
-
Maintaining world consistency when agents are occluded or camera views change abruptly, as demonstrated by successful performance on dynamic 3D Block World tasks under local-action conditioning.
-
Enabling video generation systems to learn and correct their internal state based on future observations during the inference process (Test-Time Training), leading to more robust control policies in interactive gaming or robotics.
-
Solving complex permutation problems (like S5 state tracking) where the next correct step depends on a non-trivial composition of previous actions, rather than just learning short sequences.
In summary, the improved AI systems will move beyond visual plausibility
to achieve true world understanding
by explicitly learning and maintaining a persistent, unobserved internal model of the environment's state.
Abstract
Video world models are increasingly used as simulators, but visual fidelity alone does not show that a model maintains the hidden state of the world. We examine this difference with an action-conditioned video Shell Game, a visual analogue of S 5 state tracking that separates visual rendering from compositing the unobserved world state. Trained on 5-swap chains, standard backbones (e.g., bidirectional and autoregressive Transformers, Mamba, and linear attention) render plausible videos and predict the correct ball location up to 5 swaps. However, they fail to learn the rule and generalize to longer swap chains, even with more denoising steps. As the pixel-based diffusion loss does not force the generated frames to hold the unseen ball position, output tokens cannot carry it, and the state has to live within the architecture. In a causal Transformer, this implicit state is an append-only KV cache, which is written once and never revised, so the model must re-compose the swaps at every chunk. Tracking S 5 this way requires depth to grow with sequence length, which no fixed-depth Transformer provides. We study what enables learning the rule, and find that length generalization requires a revisable state carried across chunks and an update expressive enough to apply a swap. Linear attention can achieve this by allowing negative transition eigenvalues, and autoregressive Transformers can do so with nonlinear TTT fast weights (e.g., SwiGLU) whose online updates change the feature map used to read their state. We further examine Memory Maze and Block World, where the state is not fixed by the input action stream alone and must be corrected from observations or keeps changing out of view, and discuss the implications for building stateful video world models.
Sources
- Cosmos World Foundation Model Platform for Physical AI
- MotionCraft: Physics-based Zero-Shot Video Generation
- Titans: Learning to Memorize at Test Time
- Genie: Generative Interactive Environments
- Diffusion Forcing: Next-token Prediction Meets Full-Sequence Diffusion
- Recurrent Autoregressive Diffusion: Global Memory Meets Local Attention
- Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality
- Universal Transformers
- Flex Attention: A Programming Model for Generating Optimized Attention Kernels
- DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos
- Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach
- Looped Transformers as Programmable Computers
- Force Prompting: Video Generation Models Can Learn and Generalize Physics-based Control Signals
- Unlocking State-Tracking in Linear RNNs Through Negative Eigenvalues
- Matrix-game 2.0: An open-source, real-time, and streaming interactive world model
- Denoising Diffusion Probabilistic Models
- RELIC: Interactive Video World Model with Long-Horizon Memory
- Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion
- Less is More: Recursive Reasoning with Tiny Networks
- How Far is Video Generation from World Model: A Physical Law Perspective
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models