MMGR: Multi-Modal Generative Reasoning Benchmark and Evaluation
cs.CL, cs.CV
Submitted: 2025-12-16
Updated: 2026-09-10
Comments: EMNLP 2026
Code: https://github.com/Zefan-Cai/MMGR
Project page: https://zefan-cai.github.io/MMGR.github.io
License: http://creativecommons.org/licenses/by/4.0/
The gist: Modern multimodal generative models can synthesize visually compelling images and videos, but it remains unclear whether this visual fluency reflects genuine reasoning: when prompted to generate a
Terminology
Abstract
Modern multimodal generative models can synthesize visually compelling images and videos, but it remains unclear whether this visual fluency reflects genuine reasoning: when prompted to generate a solution, can a model preserve the physical, logical, spatial, and temporal constraints a task requires, or does it merely produce plausible-looking media? To answer this question, we introduce MMGR (Multi-Modal Generative Reasoning Benchmark and Evaluation), a benchmark for evaluating generative reasoning across video, image, and language-based systems. MMGR covers 10 tasks from three domains (Abstract Reasoning, Embodied Navigation, and Physical Commonsense) and probes five reasoning abilities: Physical, Logical, 2D Spatial, 3D Spatial, and Temporal. Its evaluation emphasizes answer-verifiable tasks and, for video generation, process-aware chain-of-frame reasoning, where intermediate frames must form valid steps toward the target outcome rather than visually smooth but incorrect transitions. Evaluating state-of-the-art video generators, image generators, and LLM/VLM baselines reveals a sharp gap between visual quality and reasoning correctness: video models perform best on Physical Commonsense, but remain weak on symbolic tasks such as Sudoku, ARC, and Math, and brittle in cross-view embodied navigation. Image generators often outperform video generators on embodied navigation despite lacking temporal outputs, showing that longer visual generation does not automatically yield stronger reasoning. MMGR reframes evaluation of multimodal generation from whether outputs look realistic to whether they solve the underlying reasoning problem.
Sources
- BEVBert: Multimodal Map Pre-training for Language-guided Navigation
- VideoPhy: Evaluating Physical Commonsense for Video Generation
- VideoPhy-2: A Challenging Action-Centric Physical Commonsense Evaluation in Video Generation
- ObjectNav Revisited: On Evaluation of Embodied Agents Navigating to Objects
- From Preferences to Prejudice: The Role of Alignment Tuning in Shaping Social Bias in Video Diffusion Models
- On the Measure of Intelligence
- Training Verifiers to Solve Math Word Problems
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models
- Are Video Models Ready as Zero-Shot Reasoners? An Empirical Study with the MME-CoF Benchmark
- World Models
- Measuring Mathematical Problem Solving With the MATH Dataset
- Imagen Video: High Definition Video Generation with Diffusion Models
- Video Diffusion Models
- CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers
- Why Vision Language Models Struggle with Visual Arithmetic? Towards Enhanced Chart and Geometry Understanding
- SASRA: Semantically-aware Spatio-temporal Reasoning Agent for Vision-and-Language Navigation in Continuous Environments
- A Configurable Library for Generating and Manipulating Maze Datasets
- NavigateDiff: Visual Predictors are Zero-Shot Navigation Assistants
- Qwen-Image Technical Report
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering