Multi-Agent Egocentric World Model with Fine-Grained Embodied Interaction
cs.CV, cs.AI
Submitted: 2026-10-08
Updated: 2026-10-08
Project page: https://cvlab-kaist.github.io/ME-World
Terminology
Sources
- World Simulation with Video Foundation Models for Physical AI
- MASS: Multiplayer World Models with Authoritative Shared State
- Rhythm: Learning Interactive Whole-Body Control for Dual Humanoids
- HandsOnWorld: Unconstrained Egocentric Video Generation with Camera-Disentangled Hand Control
- DreamX-World 1.0: A General-Purpose Interactive World Model
- Vista: A Generalizable Driving World Model with High Fidelity and Versatile Controllability
- YOLOX: Exceeding YOLO Series in 2021
- Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives
- E$^3$C: Video Generation with 3D Environmental Memory and Ego-Exo Human Pose Control
- EgoSim: Egocentric World Simulator for Embodied Interaction Generation
- Multiplayer Interactive World Models with Representation Autoencoders
- MetaWorld: Scaling Multi-Agent Video World Model from Single-view Video Data
- AnchorWorld: Embodied Egocentric World Simulation with View-based Evolution Customization
- Depth Anything 3: Recovering the Visual Space from Any Views
- Gamma-World: Generative Multi-Agent World Modeling Beyond Two Players
- Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow
- Decoupled Weight Decay Regularization
- Streaming Multi-Agent Autoregressive Diffusion Model with World State Registers
- Cosmos World Foundation Model Platform for Physical AI
- MultiGen: Level-Design for Editable Multiplayer Worlds in Diffusion Game Engines
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models