Diagram-MMU: A Multi-Modal Benchmark for Scientific Diagrams
Weihao Bo, Shan Zhang, Yanpeng Sun, Jie Liu, Yongke Yao, Jinhao Du, Wei He, Kai Zou, Zechao Li, Jingdong Wang
Nanjing University of Science and Technology · Baidu Inc · Adelaide University · Singapore University of Technology and Design · Southeast University · East China Normal University · University of Oxford · NetMind.ai
cs.CV, cs.AI
Submitted: 2026-08-12
Updated: 2026-08-13
Project page: https://vi-ocean.github.io/projects/diagram-mmu
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 95/100
The gist: Diagram-MMU is a multi-modal benchmark designed to assess Multimodal Large Language Models' (MLLMs) ability for scientific diagram parsing and understanding.
Terminology
Summary
Diagram-MMU is a multi-modal benchmark designed to assess Multimodal Large Language Models' (MLLMs) ability for scientific diagram parsing and understanding. It features 3.7k curated diagrams and 18.3k human-validated questions across six domains. It evaluates MLLMs on three tasks common in vibe writing workspaces: diagram-to-code parsing, diagram-to-code editing, and diagram question answering, alongside agentic settings per task. The evaluation of 12 MLLMs reveals that diagram-to-code tasks are more challenging than diagram question answering: models can reason well over diagrams but struggle to parse and edit them, underscoring the need for methods to enhance MLLMs' capability in diagram-to-code generation. Under agentic settings, most models improve parsing and editing performance but degrade on question answering, while Claude-4.6 Opus consistently improves across all three tasks.
The benchmark comprises a carefully curated dataset sourced from TikZ/LATEX documentations, with a selection and filtering process to control quality. The final 3,744 unique diagrams paired with 18,305 evaluation samples are cross-validated by 13 graduate students. Specifically, (1) we adopt TikZ code, which is commonly used in paper writing in LATEX and integrates directly into Overleaf and Prism, and cover six scientific diagram types including charts, planar geometry, 3D shapes, graphs, chemistry, and circuit diagrams (Fig. 1)–for the first time covering chemistry and circuits for diagram-to-code tasks; (2) for evaluation metrics on diagram-to-code tasks, we use three levels: image-level measuring visual appearance similarity, code-level measuring syntactic correctness, and our object-level F1 scores measuring the basic objects that the code draws; (3) we design 16 controllable evaluation settings (Table 4): 3 for foundational ability and 13 for agentic ability–context utilization, tool use, state management, and planning. Importantly, for tool use we build our TikZ search tool as an MCP server [27], enabling MLLMs to selectively access relevant references, reducing the noise of web search and the context rot caused by loading full PDF manuals.
Overall, Diagram-MMU is the first benchmark to test both foundational and agentic abilities for solving scientific diagram-related tasks. Evaluation on 12 MLLMs (6 closed-source, 6 open-source) reveals the key findings on foundational ability: (1) Models can reason well over diagrams (DQA accuracy up to 86%) but struggle to code them (D2C-P object-level F1 ranges 31–57%), revealing perception and coding limitations; (2) Models that fail on diagram-to-code parsing also struggle with editing (D2C-E) and answering tasks (DQA); (3) Gemini-3.0 Pro achieves the most balanced profile across three tasks. On the agentic ability, (4) agency benefits editing more than parsing, likely because textual editing instructions help guide tool and context use; (5) most models degrade from excessive retrieval or poorly targeted queries in the TikZ search process, while Claude-4.6 Opus has the strongest tool use ability; (6) planning is the weakest agentic capability, especially on DQA (−0.3 to −8.8 accuracy degradation). Diagram-MMU serves as a pilot evaluation testbed for scientific diagram parsing, editing, and understanding in vibe writing workspaces, and we hope it inspires further works.
Diagram-MMU covers six scientific diagram types, each associated with one or more dedicated TikZ packages. The full dataset contains 3,744 diagrams paired with 18,305 evaluation instances (1 D2C-P + 2 D2C-E + 2 DQA per diagram), with a balanced mini split of 300 diagrams (50 per domain). The six diagram types are: Charts (960 diagrams, rendered primarily with pgfplots), Planar Geometry (601 diagrams, drawn with tikz and tkz-euclide), 3D Shapes (237 diagrams, rendered with pgfplots and tikz), Graph Structures (1,356 diagrams, drawn with tikz and tikz-network), Chemistry (187 diagrams, rendered with chemfig), and Circuit Diagrams (403 diagrams, drawn with circuitikz).
For the D2C-E task, each sample is annotated with one of four editing dimensions: text (label/annotation modification), color (fill/stroke color change), scope (local element addition, deletion, or transformation), and layout (global structural change such as chart type conversion or circuit topology modification). Layout editing is excluded from 3D shapes and chemistry. The DQA task comprises two question types: descriptive and reasoning. Reasoning questions are further split into standard and what-if types. The DQA uses 60 manually designed question templates across the six diagram domains, including descriptive questions (23 templates) that assess the model's ability to identify domain-specific symbols and extract information; standard reasoning questions (18 templates) that require numerical computation with domain-specific formulas; and what-if reasoning questions (19 templates) that require predicting answers conditioned on hypothetical element modifications.
For evaluation metrics on diagram-to-code tasks, three levels are used: image-based (SSIM, CLIP Score, LPIPS, FID), code-based (CrystalBLEU), and object-based (F1 scores for type, text, color, and bbox). For the editing task, the code- and object-based metrics are further split into preserve-only (computed over elements that should stay the same) and edit-only (computed over elements that the instruction asks to change). Diagram question answering is evaluated using accuracy, with an LLM judge (Qwen3-Next-80B-A3B-Instruct) used to extract the answer and assign binary scores.
The evaluation scheme provides 16 flexible control settings evaluating both foundational and agentic capability of current MLLMs. The four levels of agentic ability are defined as: (1) Context utilization: whether models can leverage task-relevant information from context; (2) Tool use: knowing when to invoke tools, what to query, and how to incorporate the results; (3) State management: whether models can build on prior outputs incrementally toward the final goal; (4) Planning: deciding which ability is needed and how to combine them for task solving. For tool use evaluation, a TikZ search tool is built as an MCP server, enabling selective access to relevant references instead of ingesting entire documents.
The main results on foundational capability show that: (1) A larger performance gap among open-source models in object perception exists; (2) Models follow editing instructions but struggle to produce matching code; (3) Models failing on D2C-P also struggle with D2C-E and DQA; (4) Models struggle with fine-grained spatial grounding; (5) Models struggle most with 3D shapes across all three tasks. On agentic capacity, the findings include: (1) Context utilization: most models can use objects to ground local edits but struggle to integrate them into syntax or domain knowledge reasoning; (2) Tool use: only Claude-4.6 Opus consistently benefits from tool access, while Gemini-3.1 Pro suffers from excessive retrieval and Qwen3-VL-8B fails at querying; (3) State management: most models fail to manage intermediate code states across steps, with only Claude-4.6 Opus maintaining coherence from coding to editing and reasoning; (4) Planning: planning is the weakest agentic capability, with no model reliably composing multiple abilities as task complexity grows from D2C-P to DQA.
Improvements for AI systems
Improvements to AI Systems Based on Diagram-MMU Findings:
-
Add a dedicated diagram-to-code generation module that converts visual diagrams into executable TikZ/LATEX code, rather than relying on general vision-language reasoning. This module should include object-level parsing (type, text, color, bounding box) and syntax-aware code generation, targeting the observed 31–57% object-level F1 gap.
-
Implement a hierarchical perception pipeline that separates high-level diagram understanding (for question answering) from fine-grained spatial grounding (for code generation). This addresses the finding that models reason well (up to 86% DQA accuracy) but fail at precise object localization and coding, especially for 3D shapes.
-
Integrate a selective tool-use mechanism with a TikZ reference search server (e.g., via MCP) that retrieves only relevant package documentation snippets on demand, avoiding full-manual context rot. The system should learn when to query, what to query, and how to filter results—mimicking Claude-4.6 Opus’s behavior—to prevent performance degradation from excessive or poorly targeted retrieval.
-
Add a state-management controller that tracks intermediate code states across multi-step tasks (parsing → editing → answering). This controller should maintain coherence between generated code, applied edits, and subsequent reasoning, addressing the finding that most models fail to build on prior outputs incrementally.
-
Develop a planning module that dynamically selects and sequences abilities (perception, coding, tool use, reasoning) based on task complexity. Since planning is the weakest capability (causing up to −8.8% accuracy degradation on DQA), the system should explicitly decompose tasks into sub-goals and verify each step before proceeding.
-
Enhance editing capability with dimension-aware instruction following—separate handling for text, color, scope, and layout edits. The system should preserve unchanged elements (via preserve-only metrics) while precisely modifying target elements (via edit-only metrics), using object-level F1 as a training signal to improve code-output alignment.
-
Add a domain-specific error-correction layer for diagram types where models perform worst (3D shapes, chemistry, circuits). This layer should validate generated code against known TikZ package syntax and geometric constraints, and trigger re-generation or tool queries when object-level F1 falls below a threshold.
What the Improved AI System Can Do:
-
Convert a scientific diagram (e.g., a circuit or 3D plot) into correct, compilable TikZ code with high object-level fidelity (targeting >70% F1, up from 57%).
-
Answer descriptive and reasoning questions about diagrams with >90% accuracy while simultaneously producing editable code, without performance trade-offs.
-
Perform targeted edits (e.g., change a label color or convert a chart type) while preserving all other elements, verified by both code and image-level metrics.
-
Use a TikZ search tool intelligently—querying only when needed, retrieving minimal relevant snippets, and integrating them without context overload.
-
Maintain a coherent workflow across parsing, editing, and reasoning tasks, tracking intermediate states and planning steps to avoid compounding errors.
-
Adapt to all six diagram domains (charts, geometry, 3D, graphs, chemistry, circuits) with specialized handling for the most challenging types, reducing the observed 3D-shape performance drop.
Abstract
Multimodal Large Language Models (MLLMs) have been growing the capability for scientific writing and collaboration. For example, OpenAI Prism is a free workspace for scientific writing and collaboration. One important feature in Prism is turning scientific diagrams directly into LaTeX TikZ code. In this paper, we build a benchmark, Diagram-MMU, a multi-modal benchmark designed to assess MLLMs' ability for scientific diagram parsing and understanding. Diagram-MMU features 3.7k curated diagrams and 18.3k human-validated questions across six domains. It evaluates MLLMs on three tasks common in vibe writing workspaces: diagram-to-code parsing, diagram-to-code editing, and diagram question answering, alongside agentic settings per task. The evaluation of 12 MLLMs reveals that diagram-to-code tasks are more challenging than diagram question answering: models can reason well over diagrams but struggle to parse and edit them, underscoring the need for methods to enhance MLLMs' capability in diagram-to-code generation. Under agentic settings, most models improve parsing and editing performance but degrade on question answering, while Claude-4.6 Opus consistently improves across all three tasks. Project Page: https://vi-ocean.github.io/projects/diagram-mmu.
Sources
- Kimi K2: Open Agentic Intelligence
- GLM-5: from Vibe Coding to Agentic Engineering
- Enhancing Descriptive Captions with Visual Attributes for Multimodal Perception
- Agentic Learner with Grow-and-Refine Multimodal Semantic Memory
- Math Blind: Failures in Diagram Understanding Undermine Reasoning in MLLMs
- VisCodex: Unified Multimodal Code Generation via Merging Vision and Coding Models
- ChartE$^{3}$: A Comprehensive Benchmark for End-to-End Chart Editing
- WorldVQA: Measuring Atomic World Knowledge in Multimodal Large Language Models
- Agentic Artificial Intelligence (AI): Architectures, Taxonomies, and Evaluation of Large Language Model Agents
- Human-Centric Foundation Models: Perception, Generation and Agentic Modeling
- BabyVision: Visual Reasoning Beyond Language
- Qwen3-VL Technical Report
- Kimi K2.5: Visual Agentic Intelligence
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
- Nodes Are Early, Edges Are Late: Probing Diagram Representations in Large Vision-Language Models
- Overcoming Vision Language Model Challenges in Diagram Understanding: A Proof-of-Concept with XML-Driven Large Language Models Solutions
- Visual Autoregressive Modeling for Instruction-Guided Image Editing
- TikZilla: Scaling Text-to-TikZ with High-Quality Data and Reinforcement Learning
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models