MMArch: Benchmarking Multimodal Reasoning Grounded in Architectural Evidence

arXiv:2608.09281 · cs.AI · Submitted 2026-08-10 · Read on arXiv

Chenxu Du, Kang An, Tengyue Wang, Zhongyu Yang, Xinqi Yang, Yuanchi Zhu, Hebao Zhu, Ziliang Wang, Faqiang Qian, Yunli Yang, Qibing Ren

Southwest Jiaotong University · Shanghai Jiao Tong University · ModelBest · SenseTime · South China University of Technology · East China Normal University · ShanghaiTech University · Institute of Automation, Chinese Academy of Sciences · Chongqing University · Institute for Advanced Algorithms Research, Shanghai

cs.AI

Submitted: 2026-08-10

Updated: 2026-08-11

Project page: https://dcx-swjtu.github.io/MMArch

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: MMArch: Benchmarking Multimodal Reasoning Grounded in Architectural Evidence introduces a benchmark for evaluating multimodal large language models (MLLMs) on principle-grounded visual reasoning in

Terminology

Summary

MMArch: Benchmarking Multimodal Reasoning Grounded in Architectural Evidence introduces a benchmark for evaluating multimodal large language models (MLLMs) on principle-grounded visual reasoning in architecture and civil engineering. The benchmark is built entirely from figures in peer-reviewed papers, spanning ten subdomains: structural engineering, seismic engineering, building physics, building inspection, heritage conservation, spatial analysis, building-envelope engineering, urban planning, BIM and digital construction, and computational design.

The benchmark contains 1,212 short-answer items, each requiring models to perceive relevant visual evidence across one to three figures, identify the governing engineering principle, and apply it to derive a conclusion. The construction pipeline decouples answer generation from question construction using a planner–writer approach: a planner selects an evidence lead and freezes an answer of at most ten tokens, then a writer composes a free-form question around that frozen answer. Quality control includes automated screening against text-only and caption-only shortcuts, a blind three-agent adversarial audit, dual-path answer verification, and unanimous expert review by three professional architects and engineers.

Evaluation of 18 open-weight and proprietary MLLMs against a domain-expert panel reveals a wide gap: the strongest open-source model attains about 30% and the best proprietary system 52%, while human experts reach 95%, more than forty points ahead. Specifically, GPT-5.5 achieves 51.73% and Claude Opus 4.8 achieves 51.12%, while the strongest open-source model, MiniMax M3, reaches 29.94%. The human panel averages 94.57% accuracy, with subdomain scores ranging from 89.25% to 100%.

Error analysis categorizes failures into five types: perception errors (20.3%), principle errors (21.7%), composition errors (35.8%), grounding errors (14.5%), and consistency errors (7.7%). Composition errors—where intermediate quantities are extracted correctly but the final multi-step combination mishandles a term—are the dominant failure mode, indicating that combining evidence and principle, rather than locating it, is the primary bottleneck.

Prompting interventions show limited effect: chain-of-thought reasoning yields a modest gain for GPT-4o (+1.5 points) but degrades performance for other models, and panel-identifier visual prompting produces small gains for some models while degrading others. The paper concludes that the gap points to advances in model architecture or training rather than prompting, and positions MMArch as a diagnostic testbed for such progress.

Improvements for AI systems

Improvements to AI Systems:

  1. Add a dedicated multi-step compositional reasoning module that explicitly tracks intermediate quantities and their relationships (e.g., via a symbolic scratchpad or graph-based memory) before producing a final answer. This directly targets the dominant 35.8% composition-error rate, where models extract correct evidence but fail to combine terms correctly.

  2. Train on principle-labeled visual data using a curriculum that first teaches perception of single figures, then principle identification, then multi-figure integration. The benchmark’s decoupled planner–writer pipeline can be reused as a data-generation tool to create millions of synthetic training pairs with frozen answers and varied questions, forcing the model to learn principle application rather than pattern matching.

  3. Implement a two-stage inference pipeline: Stage 1 uses a perception-only model to extract all visual elements (dimensions, labels, annotations) from each figure; Stage 2 uses a reasoning-only model that receives these structured extractions plus the question, with a hard constraint to output an answer ≤10 tokens. This separation prevents perception noise from corrupting reasoning and reduces perception errors (20.3%) and grounding errors (14.5%).

  4. Add an explicit consistency-check layer that cross-verifies the final answer against each figure’s extracted evidence and the stated principle, flagging and re-running reasoning if the answer contradicts any single figure’s data. This addresses consistency errors (7.7%) and grounding errors (14.5%) by forcing evidence alignment.

  5. Introduce a principle-retrieval pretraining task where the model is trained to map a question to the governing engineering principle (e.g., load path, thermal bridging) before seeing figures. This reduces principle errors (21.7%) by making principle identification a separate, learnable skill rather than an implicit part of end-to-end reasoning.

  6. Replace chain-of-thought prompting with a structured evidence–principle–combination template that the model must fill in, where each step is validated against a rule-based checker (e.g., units match, quantities are used exactly once). This avoids the degradation seen with CoT in other models and forces explicit composition.

  7. Add a multimodal attention-differentiation mechanism that learns to weight figures by relevance to the question, trained on the benchmark’s one-to-three figure pairings. This improves perception of cross-figure evidence and reduces composition errors by focusing computation on the most informative visual regions.

What the improved AI system can do:

  • Correctly answer principle-grounded visual reasoning tasks in architecture and civil engineering at near-expert level (target >90% accuracy on MMArch), by reliably extracting evidence from multiple figures, identifying the governing principle, and combining intermediate quantities without losing terms.

  • Generalize to other multimodal reasoning domains (e.g., medical imaging, mechanical diagrams) where multi-step evidence combination is required, because the modular architecture (perception, principle retrieval, composition, consistency check) is domain-agnostic.

  • Provide explainable outputs: for each answer, it can output the extracted evidence, the identified principle, and the step-by-step combination, enabling auditability and error correction in real-world applications like structural safety review or building code compliance.

  • Operate robustly without prompt engineering, since the structured inference pipeline replaces fragile prompting strategies, making it suitable for deployment in automated design review or educational tools.

Abstract

Multimodal large language models (MLLMs) perform strongly on engineering imagery, yet existing benchmarks mostly test drawing recognition, information extraction, or compliance checking, leaving open whether models can combine distributed visual evidence with engineering principles to reach a conclusion. We introduce MMArch, a benchmark for architecture and civil engineering spanning ten subdomains and built entirely from figures in peer-reviewed papers. Its 1, 212 short-answer items are produced by a decoupled planner--writer pipeline and validated through automated screening, a blind adversarial audit, and expert review, so that answering requires perceiving the relevant evidence, identifying the governing principle, and applying it, not exploiting textual or single-figure shortcuts. Evaluating 18 open-weight and proprietary MLLMs against a domain-expert panel, we find a wide gap: the strongest open-source model attains about 30% and the best proprietary system 52%, while human experts reach 95%, more than forty points ahead. Our error analysis shows that failures concentrate in applying principles and combining evidence across figures rather than in locating it, pointing to substantial headroom for future research. Code and data are available at https://dcx-swjtu.github.io/MMArch/.

Sources

Related papers