Modality Maturity Index: A benchmark for assessing multimodal capabilities of omni models
cs.CV, cs.AI, cs.MM
Submitted: 2026-08-26
Updated: 2026-08-26
Comments: 26 pages, 6 figures. Code and dataset available
Code: https://github.com/facebookresearch/modality-maturity-index
License: http://creativecommons.org/licenses/by/4.0/
The gist: Frontier language models are increasingly marketed as omni systems that can perceive and respond across modalities.
Terminology
Abstract
Frontier language models are increasingly marketed as omni systems that can perceive and respond across modalities. Existing evaluation frameworks, however, focus almost exclusively on bimodal understanding, typically text plus one other modality. We propose the Modality Maturity Index (MMI), a benchmark designed to evaluate the multimodal capabilities of large language models across five modalities (text, image, audio, video and document) and combinations of up to three modalities in both inputs and outputs. MMI consists of 893 questions, each carefully crafted to require the model to demonstrate its understanding of multiple input modalities and to generate responses that incorporate various output formats. The questions are designed to be self-contained, with clear expectations for the correct modality or mix of modalities required for an accurate response. Every MMI prompt carries human-authored rubric criteria for each output modality expected in the response; a model's MMI Value expresses the average of the per-modality scores for each prompt. Because low scores can reflect either failure to generate a modality (lack of presence) or failure to generate correct content, we introduce also a supplementary Modality Presence Score (MPS), a per-prompt F1 over the expected output modalities. Applying MMI to five frontier multimodal models, we find that the MPS ranges from only 15.6 (Claude Opus 4.6) to 34.9 (GPT-5.4). Given the low availability of returned modalities to even grade, we report MPS as our main result pending model improvements. To assess the viability of judging output correctness with LLM judges and rubrics, we run a separate experiment with custom generation tools. On the assets that generates, we find that an LLM judge applying the rubrics agrees with rubric-blind human annotators (who score the outputs directly and never see the criteria) on 70.8% of judgments.
Sources
- Chameleon: Mixed-Modal Early-Fusion Foundation Models
- WearVQA: A Visual Question Answering Benchmark for Wearables in Egocentric Authentic Real-world scenarios
- Are We on the Right Way for Evaluating Large Vision-Language Models?
- MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
- MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
- Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
- BLINK: Multimodal Large Language Models Can See but Not Perceive
- WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs
- MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective
- T2I-CompBench++: An Enhanced and Comprehensive Benchmark for Compositional Text-to-image Generation
- GenAI-Bench: Evaluating and Improving Compositional Text-to-Visual Generation
- SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
- LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods
- MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
- OmniBench: Towards The Future of Universal Omni-Language Models
- MMBench: Is Your Multi-modal Model an All-around Player?
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
- Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration
- MMLongBench-Doc: Benchmarking Long-context Document Understanding with Visualizations
- EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models