minWM: A Full-Stack Open-Source Framework for Real-Time Interactive Video World Models
cs.CV
Submitted: 2026-05-28
Updated: 2026-09-22
Code: https://github.com/shengshu-ai/minWM
Terminology
Sources
- Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models
- CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer
- Open-Sora Plan: Open-Source Large Video Generation Model
- Open-Sora: Democratizing Efficient Video Production for All
- Wan: Open and Advanced Large-Scale Video Generative Models
- HunyuanVideo: A Systematic Framework For Large Video Generative Models
- WorldPlay: Towards Long-Term Geometric Consistency for Real-Time Interactive World Modeling
- Hunyuan-GameCraft-2: Instruction-following Interactive Game World Model
- Yume-1.5: A Text-Controlled Interactive World Generation Model
- Vidarc: Embodied Video Diffusion Model for Closed-loop Control
- Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length
- StreamAvatar: Streaming Diffusion Models for Real-Time Interactive Human Avatars
- RELIC: Interactive Video World Model with Long-Horizon Memory
- Yan: Foundational Interactive Video Generation
- PAN: A World Model for General, Actionable, and Long-Horizon World Simulation
- Matrix-game 2.0: An open-source, real-time, and streaming interactive world model
- MotionStream: Real-Time Video Generation with Interactive Motion Controls
- Diffusion Adversarial Post-Training for One-Step Video Generation
- Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion
- Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video Generation
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models