UReason: Benchmarking Reasoning-to-Generation Alignment in Unified Multimodal Models
cs.CL, cs.CV
Submitted: 2026-02-09
Updated: 2026-08-30
Comments: Project page: https://ureason.github.io
Project page: https://ureason.github.io
License: http://creativecommons.org/licenses/by/4.0/
The gist: Unified multimodal models (UMMs) aim to integrate multimodal understanding and generation within a unified architecture, yet it remains unclear to what extent textual and visual modalities are
Terminology
Abstract
Unified multimodal models (UMMs) aim to integrate multimodal understanding and generation within a unified architecture, yet it remains unclear to what extent textual and visual modalities are aligned. To investigate this question, we use reasoning-guided image generation as a diagnostic task, where models produce textual reasoning first and then generate images. We introduce UReason, a benchmark for evaluating reasoning-to-generation alignment in this paradigm, consisting of 2,000 human-curated and human-verified instances spanning five reasoning-intensive tasks: Code, Arithmetic, Spatial, Attribute, and Text. To enable controlled analysis, we develop an evaluation framework that compares direct generation, reasoning-guided generation, and decontextualized generation, which conditions only on the refined prompt extracted from reasoning. Across eight widely used open-source UMMs, while we find that reasoning-guided generation yields improvements over direct generation, somewhat surprisingly, decontextualized generation consistently outperforms reasoning-guided generation by a large margin. Our further analyses suggest that the intended visual semantics in textual reasoning are not reliably reflected in the generated images, despite their unified design and training. Overall, UReason serves as a practical litmus test for reasoning-to-generation alignment and provides a challenging benchmark for developing next-generation, more tightly aligned UMMs.
Sources
- R2I-Bench: Benchmarking Reasoning-Driven Text-to-Image Generation
- Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models
- Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling
- Emerging Properties in Unified Multimodal Pretraining
- Commonsense-T2I Challenge: Can Text-to-Image Generation Models Understand Commonsense?
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Seed1.5-VL Technical Report
- Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models
- T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT
- SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models
- Uni-MoE-2.0-Omni: Scaling Language-Centric Omnimodal Large Model with Advanced MoE, Training and Data
- Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models
- Does Understanding Inform Generation in Unified Multimodal Models? From Analysis to Path Forward
- WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation
- Muddit: Liberating Generation Beyond Text-to-Image with a Unified Discrete Diffusion Model
- T2I-ReasonBench: Benchmarking Reasoning-Informed Text-to-Image Generation
- Unified Multimodal Discrete Diffusion
- Chameleon: Mixed-Modal Early-Fusion Foundation Models
- Beyond Language Modeling: An Exploration of Multimodal Pretraining
- Emu3: Next-Token Prediction is All You Need
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering