Zing-0.5: Toward Playable Worlds with Real-Time Joint Action and Text Control
cs.CV, cs.LG
Submitted: 2026-09-15
Updated: 2026-09-15
Comments: 19 pages, 8 figures. Authors listed alphabetically by surname. Project: https://zing.loopit.me/ ; Code: https://github.com/seedleap/zing-world-model ; Models: https://huggingface.co/seedleap/zing-0.5 ; Serving: https://github.com/seedleap/Zing-SGLang
Code: https://github.com/seedleap/zing-world-model
Project page: https://meituan-longcat.github.io
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- GameGen-X: Interactive Open-world Game Video Generation
- Diffusion Forcing: Next-token Prediction Meets Full-Sequence Diffusion
- H3-World: Turning Language Understanding into World Control
- Data-Forcing Distillation: Restoring Diversity and Fidelity in Few-Step Video Generation
- LongLive-2.0: An NVFP4 Parallel Infrastructure for Long Video Generation
- Infinite Worlds with Versatile Interactions
- CameraCtrl: Enabling Camera Control for Text-to-Video Generation
- Matrix-game 2.0: An open-source, real-time, and streaming interactive world model
- Classifier-Free Diffusion Guidance
- CMT: Mid-Training for Efficient Learning of Consistency, Mean Flow, and Flow Map Models
- DisCo: World Models with Discrete Camera Motion Control
- Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion
- ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU
- Hunyuan-GameCraft: High-dynamic Interactive Game Video Generation with Hybrid History Condition
- Flow Matching for Generative Modeling
- Decoupled DMD: CFG Augmentation as the Spear, Distribution Matching as the Shield
- CausalCine: Real-Time Autoregressive Generation for Multi-Shot Video Narratives
- WorldCam: Interactive Autoregressive 3D Gaming Worlds with Camera Pose as a Unifying Geometric Representation
- Advancing Open-source World Models
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models