H3-World: Turning Language Understanding into World Control
cs.CV, cs.AI
Submitted: 2026-09-01
Updated: 2026-09-01
Code: https://github.com/Danzer1xxxxChan/H3-World
Project page: https://danzer1xxxxchan.github.io/H3-World
License: http://creativecommons.org/licenses/by/4.0/
The gist: We present H3-World, an efficient framework that turns the 33B MiniMax-H3 video generator into an interactive world model.
Terminology
Abstract
We present H3-World, an efficient framework that turns the 33B MiniMax-H3 video generator into an interactive world model. Our key finding is that, as large video generators become more capable, language is emerging as a natural interface for control. MiniMax-H3, for example, already supports zero-shot control of character behavior and camera motion through natural-language instructions. Building on this, H3-World turns this coarse language interface into precise, temporally grounded world control, without introducing dedicated action modules. Specifically, we represent each action as a structured combination of character and camera instructions, and align them with the corresponding temporal video latents. To make the control temporally precise, we further introduce temporal attention routing, which restricts each instruction to its intended time interval and reduces control leakage across actions. Importantly, H3-World directly reuses the semantic representations learned during large-scale video pretraining and requires only lightweight adaptation. With only 8,000 gameplay samples, 10,000 LoRA optimization steps, and 0.199% trainable parameters, H3-World achieves effective character and camera control while preserving strong generation quality. It also generalizes to unseen scenarios. These results show that the control capabilities emerging in large video generators can be efficiently transformed into interactive world control.
Sources
- Wan: Open and Advanced Large-Scale Video Generative Models
- World Models
- Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets
- Latte: Latent Diffusion Transformer for Video Generation
- Open-Sora Plan: Open-Source Large Video Generation Model
- HunyuanVideo: A Systematic Framework For Large Video Generative Models
- Learning Interactive Real-World Simulators
- Matrix-Game: Interactive World Foundation Model
- WorldPlay: Towards Long-Term Geometric Consistency for Real-Time Interactive World Modeling
- Advancing Open-source World Models
- Infinite Worlds with Versatile Interactions
- Matrix-game 2.0: An open-source, real-time, and streaming interactive world model
- Matrix-Game 3.0: Real-Time and Streaming Interactive World Model with Long-Horizon Memory
- ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU
- ReactiveGWM: Flexible Control and NPC Reactivity in Game World Models
- SCOPE: Simulating Cross-game Operations in Playable Environments for FPS World Models
- StatePlay: State-Aware Game World Models for Mechanics-Consistent Generation
- BadWAM: When World-Action Models Dream Right but Act Wrong
- BadWorld: Adversarial Attacks on World Models
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models