Better Slots, Better Worlds: Representation Quality & Robustness in Object-Centric World Models
Shukrullo Nazirjonov, Sai Prasanna, Anna Manasyan, Georg Martius
University of Tuebingen · Max Planck Institute for Intelligent Systems
cs.CV, cs.AI, cs.LG
Submitted: 2026-08-12
Updated: 2026-08-13
Comments: Published at Model-Based RL in the Era of Generative World Models Workshop at RLC 2026
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 75/100
Terminology
Summary
arXiv: 2608.12078v1 [cs.CV], 12 Aug 2026
The paper investigates whether object-centric (OC) representations actually deliver for planning in world models and what drives their performance. The authors conduct a controlled study of object-centric world models (OCWMs) for visual model-predictive control along two axes: object-centric representation quality and generalization under distribution shift relative to scene-centric models. Their key findings are:
-
Planning success correlates positively with unsupervised slot-quality metrics (FG-ARI, mBO), though gains saturate at high slot quality.
-
With well-bound slots, auxiliary proprioception inputs and masking inductive bias become unnecessary.
-
An OCWM with well-bound slots plans more robustly under unseen distribution shifts than the end-to-end trained scene-centric LeWM, while DINO-WM, built on similar frozen pretrained features, remains comparably robust — suggesting that pretrained visual representations are an important contributor to robustness.
The paper addresses a central promise of world models: that an agent can learn its environment's dynamics from offline trajectories and plan toward arbitrary goals at test time. The authors note that prior object-centric world models (OCWMs) take the slot encoder as given and evaluate only in-distribution, leaving open whether the object-centric bias actually delivers for planning and what within the OCWM drives it.
Two untested assumptions are identified:
-
Slot encoder quality: "The slot encoder is trained separately from the world model and selected using unsupervised slot-quality metrics (FG-ARI and mBO) that reward clean object masks rather than planning success; whether better-scoring slots yield better downstream planning is therefore unknown."
-
Generalization:
Whether the promised generalization holds for planning under distribution shift is untested.
The study evaluates on 2D PushT and 3D OGBench-Cube environments. The framework builds on C-JEPA (an object-centric world model), with several improvements:
-
Replacing VideoSAUR with SlotContrast as the slot encoder, which
provides stronger temporal consistency and eliminates the need for Hungarian matching across slots
-
Updating DINOv2 to DINOv3 as the feature extractor
-
The default model, SlotContrast-WM, uses the non-causal OC-JEPA backbone
with neither the slot-masking objective nor the auxiliary proprioception token
Baselines compared:
-
DINO-WM: Uses frozen DINOv2 patch tokens
-
LeWM: Uses the global CLS token of an end-to-end trained ViT
Metrics:
-
Slot quality: Video FG-ARI (Foreground Adjusted Rand Index) and video mBO (mean Best Overlap)
-
Planning: Task success rate
The authors observed that C-JEPA's VideoSAUR encoder produces poorly bound slots, fragmenting objects and bleeding them into the background,
yet C-JEPA reports strong planning performance. This raised the question: does unsupervised slot quality bear on planning success at all?
Holding the dynamics model and planner fixed, they trained SlotContrast-WM on intermediate SlotContrast checkpoints spanning a range of slot quality. Results show:
Downstream planning performance closely tracks the emergence of object separation.
On PushT, planning success rate above random baseline correlates strongly with:
-
Video FG-ARI: Pearson r = 0.96
-
Video mBO: Pearson r = 0.94
Takeaway 1: "Planning success is positively correlated with slot quality with no change to the dynamics model, but the gains saturate: at high slot quality, where the metrics themselves plateau, better-bound slots yield little additional planning success."
On OGBench-Cube, the sweep saturates early because FG-ARI and mBO are aggregated over the whole scene, where the large, easily segmented robot arm dominates the score while the small, task-relevant cube contributes little.
The authors hypothesized that the auxiliary mechanisms of masked-slot history prediction objective and the auxiliary proprioception token, which C-JEPA inherits, do not add predictive power but compensate for weak slots.
They swept num masked slots ∈ 0, 1, 2 with and without proprioception on both high-quality (SlotContrast) and low-quality (VideoSAUR) encoders. Key results on PushT:
-
SlotContrast-WM without proprioception or masking: 84.7% (±1.9) SR
-
Adding proprioception token: barely helps (+0.6 pp)
-
Full C-JEPA recipe (VideoSAUR with proprioception and masking): 85.3% (±3.4) SR
-
VideoSAUR minimal configuration: 74.7% SR (10 pp lower than SlotContrast-WM)
Critically:
"Without proprioception, masking monotonically degrades both encoders... Masking helps only when paired with proprioception on the weak encoder (73.3 → 85.3%), suggesting it leans on a proprioceptive shortcut rather than exploiting visual object interactions."
Takeaway 2: In our manipulation tasks, a sufficiently object-centric representation makes both slot history masking and proprioception unnecessary: they only compensate for weak representations.
The authors evaluated world-model robustness under visual and dynamics shifts, including object-level appearance shifts, frame-level perturbations (e.g., background color changes), scale changes, and object shape changes.
Key results on PushT:
-
SlotContrast-WM retains the highest success rates under object-level appearance shifts
-
DINO-WM degrades only moderately
-
LeWM collapses (e.g., under background color changes, LeWM drops to 2-4% SR, near random)
The authors note:
"DINO-WM plans over frozen pretrained DINO features directly, while SlotContrast-WM plans over object slots that Slot Attention extracts from the same features; the shared pretrained foundation appears to be the main source of robustness to appearance shifts."
On OGBench-Cube:
-
SlotContrast-WM and DINO-WM stay close to their in-distribution success across all shifts
-
LeWM degrades under scene-level shifts (background color and camera angle bring it close to the 48% random-policy baseline)
Notably: every model fails under geometric variations, since changing an object's shape alters its contact dynamics.
Takeaway 3: World models planning over frozen pretrained features — object-centric or not — tolerate appearance and frame-level shifts far better than the end-to-end trained LeWM.
The paper's central conclusion:
"Our controlled study isolates what makes object-centric world models work for planning: representation quality is the dominant factor — planning success closely tracks slot quality and saturates once slots are well-bound, at which point the auxiliary proprioception input and slot-masking objective become unnecessary. Given good slots, the resulting OCWM is also the most robust model in our study, though DINO-WM, built on similar frozen pretrained features, retains much of this robustness while the end-to-end trained LeWM degrades severely. In short, better slots enable stronger object-centric planning, while frozen pretrained visual representations appear to be an important ingredient for robustness under distribution shift."
"Unsupervised slot metrics are less informative when task-relevant objects are small relative to the scene (e.g., OGBench-Cube), motivating task-aware quality measures and end-to-end encoder–world-model training. Future work should test these findings in more diverse environments—more objects, varied scales, richer dynamics—to probe OCWM's robustness and compositional generalization."
Improvements for AI systems
Based on the paper, here are the specific improvements I can make to AI systems and what the improved systems can do:
-
Improvement: Replace or augment unsupervised metrics (FG-ARI, mBO) with task-relevant weighting that accounts for object size and importance relative to the planning goal. This addresses the failure mode where large, irrelevant objects (e.g., robot arm) dominate quality scores while small task-critical objects (e.g., cube) are ignored.
-
What the improved system can do: Select slot encoders based on planning-relevant object separation rather than scene-wide statistics, leading to better world models in environments with heterogeneous object sizes.
-
Improvement: Implement an adaptive mechanism that automatically disables proprioception inputs and slot-masking objectives when slot quality exceeds a learned threshold. The system can detect when slots are well-bound and switch to the minimal configuration (84.7% SR) instead of carrying unnecessary computational overhead.
-
What the improved system can do: Dynamically reduce model complexity and inference cost in deployment while maintaining peak planning performance, without requiring manual tuning per environment.
-
Improvement: Modify end-to-end trained models like LeWM to incorporate frozen pretrained visual features (e.g., DINOv3 patch tokens) as a parallel pathway alongside the trainable encoder. This bridges the robustness gap (LeWM drops to 2-4% SR under background shifts) without sacrificing end-to-end adaptability.
-
What the improved system can do: Achieve near-DINO-WM robustness under appearance and frame-level distribution shifts while retaining the flexibility of end-to-end training for task-specific dynamics learning.
-
Improvement: Add a confidence module that monitors real-time slot quality (FG-ARI/mBO) during planning and adjusts action selection—falling back to more conservative policies or requesting human intervention when slot quality degrades below a saturation threshold.
-
What the improved system can do: Gracefully handle out-of-distribution scenarios where object segmentation fails, preventing catastrophic planning failures and improving safety in real-world deployment.
-
Improvement: Incorporate multi-scale slot attention that explicitly handles objects of varying sizes by using hierarchical or multi-resolution slot extraction, addressing the OGBench-Cube limitation where small task-relevant objects are underweighted.
-
What the improved system can do: Maintain high planning success in scenes with both large manipulators and small target objects, extending OCWM applicability to more realistic robotic manipulation tasks.
-
Improvement: Implement a meta-selection mechanism that evaluates candidate world models (OCWM vs. DINO-WM vs. LeWM) based on predicted distribution-shift exposure, choosing the most robust model for the anticipated test-time conditions rather than relying solely on in-distribution performance.
-
What the improved system can do: Automatically trade off between models—using LeWM for tightly controlled environments and SlotContrast-WM or DINO-WM for deployment with unknown visual perturbations—optimizing for worst-case robustness.
Abstract
Learning world models from offline trajectories enables agents to accomplish different tasks through planning. Object-centric (OC) representations, which decompose a scene into a set of slots that bind to its objects, have been proposed as an inductive bias for world models that are more sample-efficient and generalize better. Yet prior object-centric world models (OCWMs) take the slot encoder as given and evaluate only in-distribution, leaving open whether the object-centric bias actually delivers for planning and what within the OCWM drives it. We conduct a controlled study of OCWMs for visual model-predictive control along two axes: object-centric representation quality and generalization under distribution shift relative to scene-centric models. We find that (i) planning success correlates positively with unsupervised slot-quality metrics (FG-ARI, mBO), though the gains saturate at high slot quality; (ii) with well-bound slots, the auxiliary proprioception inputs and masking inductive bias that prior methods relied on become unnecessary; and (iii) under unseen distribution shifts, the OCWM with well-bound slots is more robust overall than the end-to-end trained scene-centric LeWM, while DINO-WM, built on similar frozen pretrained features, remains competitive -- pointing to pretrained features as a key contributor to robustness.
Sources
- MONet: Unsupervised Scene Decomposition and Representation
- Learning Interactive World Model for Object-Centric Reinforcement Learning
- stable-worldmodel-v1: Reproducible World Modeling Research and Evaluation
- LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels
- Causal-JEPA: Learning World Models through Object-Level Latent Masking
- Slot-MPC: Goal-Conditioned Model Predictive Control with Object-Centric Representations
- Object-Centric World Models from Few-Shot Annotations for Sample-Efficient Reinforcement Learning
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models