Can Video World Models Track Unobserved World States?
summary
The gist
Video world models are increasingly used as simulators, yet visual fidelity alone does not show that a model maintains the hidden state of the world.
In short
The study tested if video world models can track hidden states across sequences of actions, using a 'Shell Game' as a test for state tracking. Standard models failed to maintain the hidden state when swaps exceeded a small depth. The findings suggest current architectures need recurrent transitions or history-dependent updates to robustly compose unobserved states.
Key concepts
- State Tracking
- This is the ability of a world model to remember what is not currently visible in the video and correctly update that memory as new actions and observations occur. It ensures the model can predict future frames based on past hidden information, not just what is currently seen.
- Recurrent Transition
- This refers to how a model's internal state evolves from one time step to the next. The paper suggests this needs to be expressive enough—like a nonlinear RNN—to allow the model to correctly compose complex transitions for hidden objects, rather than just adding new information linearly.
- Test-Time Training (TTT) Fast Weights
- This is a mechanism where the model's update process changes during inference. By making this update depend on its own previous state, it creates an effective kernel history. This allows the model to behave differently at test time than it did during training, helping it maintain a persistent internal state.
- Autoregressive KV Cache
- This is a memory structure used in Transformer models to store past information. The paper argues this cache is not a compact running state; instead, it forces the model to re-derive the arrangement from the entire history every time, limiting its ability to carry a simple, persistent hidden state.
Terminology used across episodes
This episode discusses
- Can Video World Models Track Unobserved World States? · Paper Radio
- Cosmos World Foundation Model Platform for Physical AI
- MotionCraft: Physics-based Zero-Shot Video Generation
- Titans: Learning to Memorize at Test Time
- Genie: Generative Interactive Environments
- Diffusion Forcing: Next-token Prediction Meets Full-Sequence Diffusion
- Recurrent Autoregressive Diffusion: Global Memory Meets Local Attention
- Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality
- Universal Transformers
- Flex Attention: A Programming Model for Generating Optimized Attention Kernels
- DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos
- Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach
- Looped Transformers as Programmable Computers
- Force Prompting: Video Generation Models Can Learn and Generalize Physics-based Control Signals
- Unlocking State-Tracking in Linear RNNs Through Negative Eigenvalues
- Matrix-game 2.0: An open-source, real-time, and streaming interactive world model · Paper Radio
- Denoising Diffusion Probabilistic Models
- RELIC: Interactive Video World Model with Long-Horizon Memory
- Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion
- Less is More: Recursive Reasoning with Tiny Networks
- How Far is Video Generation from World Model: A Physical Law Perspective
The paper
Can Video World Models Track Unobserved World States? · Read on arXiv
Joonghyuk Shin, Yicong Hong, Jaesik Park, Xun Huang
Seoul National University · Roblox
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "Can Video World Models Track Unobserved World States?".
Tom: Video world models are increasingly used as simulators, yet visual fidelity alone does not show that a model maintains the hidden state of the world.
Jane: First, who's behind it and why it matters.
Paper summary: Tom: Hey team, I’m really pumped about this paper we’re talking about today because it tackles a fundamental issue with video world models: whether they can actually remember what's hidden from view.
Jane: It sounds like the core idea is that just having good visuals isn't enough; the model needs to maintain a consistent internal understanding of the scene even when things are occluded or unseen, right?
Lu: Exactly. This paper, "Can Video World Models Track Unobserved World States?", looks at this using an action-conditioned video Shell Game as a way to test how these models handle tracking tasks like S5 state tracking.
Meng: So, if we look at the abstract from this paper, what's the main problem they're setting out to solve with this study?
Tom: Well, the paper points out that current video model backbones trained only on visible content struggle to carry unobserved states when new actions or observations arrive later on. The authors examine this gap by using an action-conditioned video Shell Game as a visual version of S5 state tracking.
Jane: So, essentially they're asking if these models can remember the ball's position even when it’s hidden across several swaps in that game?
Lu: That’s the main thesis: they show that standard architectures—like autoregressive Transformers or Mamba-based models—can fit a training horizon of five swaps, but they fail to generalize robustly when the swap chain gets longer.
Meng: So, what's the specific reason why they fail on longer chains? Is it just a limitation of the architecture itself?
Tom: Not exactly; they find that this failure happens because the pixel-based diffusion target doesn't supervise that unseen hidden state; instead, those generated frames can't carry it and the state has to live inside the architecture rather than in the tokens.
Jane: That means when you look at a Transformer, its architectural state is just an append-only KV cache, which leads to each chunk having to re-derive the arrangement from the whole history at a fixed depth.
Lu: Right, and they identify two specific mechanisms that allow models to extrapolate beyond their training length. First, they look at widening the recurrent transition for linear attention variants from zero one to
−one one: , which allows for reflection-like transitions with negative eigenvalues necessary for composing a swap.
Paper summary: Meng: And the second mechanism involves a Test-Time Training fast weight whose online nonlinear feature-map update makes its effective kernel history-dependent, moving it outside plain linear attention by rewriting the SwiGLU inner model.
Tom: So, while these mechanisms help them fit the training horizon of five swaps, they still fall toward chance on longer swap chains and provide no benefit when rendering plausible video with extra denoising steps.
Jane: It seems like even with these architectural tweaks, the pixel-based diffusion target never supervises the unseen hidden state in a way that helps the model learn to carry it forward.
Lu: The study also looked at probing internal states by saving model readouts and using fixed-bank readouts as queries into the KV cache or input to those fast weights, and they found that while visible swap identity decodes at one point zero, hidden-ball accuracy remains quite weak across the models they tested.
Meng: That suggests that we might be using these readouts mostly for rendering the current view rather than actually carrying a compact persistent state needed for tracking.
Tom: It really emphasizes that the current setup means every serial step is spent on rendering, and we observe no scratchpad emerging when you look at the rollout in DiT, unlike how chain-of-thought in language models works with working memory.
Jane: So the paper suggests that for a video model to be useful beyond short sequences, it needs a state carried across chunks paired with an update rule expressive enough to compose that hidden transition on that state in place.
Lu: That points toward needing architectures where the update isn't just additive, but history-dependent, and where the model can correct its state from observations as well.
Meng: From an engineering standpoint, if we have to re-derive the arrangement from history at every chunk for Transformers, that sounds computationally expensive when you scale up to complex scenarios.
Tom: It is expensive if you don't have a compact way to store and revise that state during the sequence generation process.
Jane: So, what does this imply for how we design video world models moving forward? Does it suggest a shift away from purely additive updates toward something more integrated?
Paper summary: Lu: The paper suggests that reliable tracking in these settings will likely require progress in stateful architectures and objectives that reach beyond just the visible content.
Meng: I wonder if this means we need to prioritize developing mechanisms like those fast weights or recurrent transitions over simply increasing the number of tokens for memory.
Tom: Precisely, because the findings show that simply adding scratchpad tokens doesn't guarantee success up to longer swap chains; LaCT performed better than plain additive updates because it normalized states via those fast weights and kept extrapolating as a better-regularized register.
Jane: So, the implication is that the focus needs to shift from just fitting the training length to creating architectures that inherently support state composition across unseen parts of a sequence.
Lu: We are seeing evidence that current AR video backbones can't compose unobserved state updates, which means they should fail whenever a correct frame depends on a long swap chain among five or more hidden objects.
Meng: If we want to deploy these in complex simulators, the challenge isn't just rendering plausible frames; it's ensuring that internal logic remains consistent over long interactions.
Tom: So, this paper gives us clear direction: we need state carried across chunks paired with an update rule expressive enough to compose the hidden transition on that state in place.
Jane: It’s a call for architectures where the update isn't merely additive but history-dependent, and where the model can correct its state from observations as well.
Lu: These findings have broad implications for dynamic world exploration tasks like Memory Maze and three dee Block World, which require a state that is not just a pure function of the action stream.
Meng: I think we need to look at how we integrate observation updates directly into the state transition mechanism, rather than treating them as just another input to the rendering pipeline.
Tom: It really highlights that while LaCT is one of the strongest backbones across most settings, achieving that correction from observations remains an open challenge for video world models.
Jane: So, we’re looking at a future where reliable tracking will likely require progress in stateful architectures and objectives that reach beyond just visible content.
Conclusion: Segment: Conclusion — The Big Picture** **(Recap)** So, to wrap up our discussion on this paper, we've seen how they tested video world models with a shell game and found that current methods struggle when the hidden state gets too complicated or long.
Tom: I think the title itself really captures the essence of the problem they are solving here—can these models actually keep track of what's happening behind closed doors in a dynamic environment? The authors did a solid job showing exactly where those models break down when it comes to memory.
Jane: It’s fascinating because it moves past just making pretty videos and gets into the core challenge of intelligence: maintaining an internal model of reality, even when you can't see everything. They show that the visual information alone isn't enough to sustain that internal knowledge over time.
Lu: From a theoretical standpoint, the authors are pushing us toward a new way of thinking about video processing, suggesting we need recurrent transitions or explicit memory structures because simple autoregressive flows just aren't structured to handle this kind of sequential state dependency effectively.
Meng: I look at the practical implications here and see that if we can't reliably track those unobserved states, then any complex simulation involving long interactions or hidden objects in a physical space will be fundamentally limited by the model’s inability to maintain a consistent world state.
Lalam: If this research moves us toward architectures that can genuinely compose hidden transitions in place, it opens up avenues for much richer cultural simulations where AI agents can learn and remember complex social or physical scenarios more deeply.
Tom: Exactly, and I think the authors' findings suggest that we need to look beyond just making the next frame look good; we need to build systems where the model is actively managing a persistent internal representation of the world.
Jane: It really boils down to this: for video models to be truly useful in dynamic exploration, they have to become stateful entities that can update their knowledge based on what they observe, not just blindly generate what looks right next.
Lu: That means the future direction involves integrating observation updates directly into the state transition mechanism so the model can correct its internal representation when it gets new input.
Meng: That correction ability is crucial; if a model drifts from its internal world state, it loses track of reality in a complex scenario, which is a major hurdle for deployment in anything beyond simple demonstrations.
Lalam: I think this work paves the way for AI systems that possess more robust "working memory," allowing them to handle long-term planning and context far better than we see now. **(Hook)** This brings us perfectly to how these concepts translate into actual design changes, which is what we'll explore next with our engineering lead.
More episodes
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language
- 2508.08833-An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
- 2405.04118-Policy Learning with a Language Bottleneck