Code World Model: Coding Agent as World Brain
cs.CV, cs.AI, cs.CL
Submitted: 2026-08-26
Updated: 2026-08-26
Comments: Project Page: https://buaacyw.github.io/cwm/
Project page: https://buaacyw.github.io/cwm
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- World Models
- Mastering Diverse Domains through World Models
- GAIA-1: A Generative World Model for Autonomous Driving
- Learning Interactive Real-World Simulators
- DreamGen: Unlocking Generalization in Robot Learning through Video World Models
- World-Gymnast: Training Robots with Reinforcement Learning in a World Model
- Pandora: Towards General World Model with Natural Language Actions and Video States
- Pre-Trained Video Generative Models as World Simulators
- Hunyuan-GameCraft: High-dynamic Interactive Game Video Generation with Hybrid History Condition
- Matrix-Game: Interactive World Foundation Model
- PAN: A World Model for General, Actionable, and Long-Horizon World Simulation
- DreamX-World 1.0: A General-Purpose Interactive World Model
- Infinite Worlds with Versatile Interactions
- WorldMem: Long-term Consistent World Simulation with Memory
- WorldPlay: Towards Long-Term Geometric Consistency for Real-Time Interactive World Modeling
- RELIC: Interactive Video World Model with Long-Horizon Memory
- Advancing Open-source World Models
- Matrix-Game 3.0: Real-Time and Streaming Interactive World Model with Long-Horizon Memory
- Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval
- VMem: Consistent Interactive Video Scene Generation with Surfel-Indexed View Memory
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models