World Observer: Joint Actor-Observer Generation for Persistent World Modeling
cs.CV
Submitted: 2026-10-01
Updated: 2026-10-01
Project page: https://cvlab-kaist.github.io/world-observer
Terminology
Sources
- SAM 3: Segment Anything with Concepts
- 360+x: A Panoptic Multi-modal Scene Understanding Dataset
- MemoBench: Benchmarking World Modeling in Dynamically Changing Environments
- Out of Sight but Not Out of Mind: Hybrid Memory for Dynamic Video World Models
- Pantheon360: Taming Digital Twin Generation via 3D-Aware 360{\deg} Video Diffusion
- DreamX-World 1.0: A General-Purpose Interactive World Model
- LiveWorld: Simulating Out-of-Sight Dynamics in Generative Video World Models
- GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium
- Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion
- VBench: Comprehensive Benchmark Suite for Video Generative Models
- PanoWorld: Geometry-Consistent Panoramic Video World Modeling
- CubeDiff: Repurposing Diffusion-Based Image Models for Panorama Generation
- HunyuanVideo: A Systematic Framework For Large Video Generative Models
- OmniNWM: Omniscient Driving Navigation World Models
- MagicWorld: Towards Long-Horizon Stability for Interactive Video World Exploration
- PanoWorld: Real-World Panoramic Generation
- CubeComposer: Spatio-Temporal Autoregressive 4K 360{\deg} Video Generation from Perspective Video
- Depth Anything 3: Recovering the Visual Space from Any Views
- Flow Matching for Generative Modeling
- OmniRoam: World Wandering via Long-Horizon Panoramic Video Generation
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models