Does Understanding Inform Generation in Unified Multimodal Models? From Analysis to Path Forward
cs.CV, cs.CL
Submitted: 2025-11-25
Updated: 2026-09-25
Code: https://github.com/vllm-project/vllm
Terminology
Sources
- The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"
- Multimodal Language Models as Text-to-Image Model Evaluators
- BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset
- R2I-Bench: Benchmarking Reasoning-Driven Text-to-Image Generation
- Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling
- Emerging Properties in Unified Multimodal Pretraining
- PUMA: Empowering Unified MLLM with Multi-granular Visual Generation
- X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again
- Turning Internal Gap into Self-Improvement: Promoting the Generation-Understanding Unification in MLLMs
- CLIPScore: A Reference-free Evaluation Metric for Image Captioning
- UniCTokens: Boosting Personalized Understanding and Generation via Unified Concept Tokens
- UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation
- UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning
- WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation
- Muddit: Liberating Generation Beyond Text-to-Image with a Unified Discrete Diffusion Model
- T2I-ReasonBench: Benchmarking Reasoning-Informed Text-to-Image Generation
- Chameleon: Mixed-Modal Early-Fusion Foundation Models
- MetaMorph: Multimodal Understanding and Generation via Instruction Tuning
- Ovis-U1 Technical Report
- Emu3: Next-Token Prediction is All You Need
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models