Can MiniMax-H3 Reason About the Physical World? An Evaluation of Omni-Modal Generative Model
cs.CV
Submitted: 2026-09-16
Updated: 2026-09-16
Code: https://github.com/gulucaptain/MiniMax-H3-Reason
Terminology
Sources
- HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone
- TiViBench: Benchmarking Think-in-Video Reasoning for Video Generative Models
- Do Joint Audio-Video Generation Models Understand Physics?
- TC-Bench: Benchmarking Temporal Compositionality in Text-to-Video and Image-to-Video Generation
- LTX-Video: Realtime Video Latent Diffusion
- Baichuan-Omni-1.5 Technical Report
- PhyGround: Benchmarking Physical Reasoning in Generative World Models
- RISE-Video: Can Video Generators Decode Implicit World Rules?
- Can World Simulators Reason? Gen-ViRe: A Generative Visual Reasoning Benchmark
- Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models
- Ola: Pushing the Frontiers of Omni-Modal Language Model
- Humanoid Manipulation Interface: Humanoid Whole-Body Manipulation from Robot-Free Demonstrations
- Seedance 2.0: Advancing Video Generation for World Complexity
- Evaluating Newtonian Mechanics in Video Generative Models with Real Physical Systems
- Wan: Open and Advanced Large-Scale Video Generative Models
- Video models are zero-shot learners and reasoners
- Omni-WorldBench: Towards a Comprehensive Interaction-Centric Evaluation for World Models
- WorldMark: A Unified Benchmark Suite for Interactive Video World Models
- Qwen3 Technical Report
- HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models