mChartQA: A universal benchmark for multimodal Chart Question Answer based on Vision-Language Alignment and Reasoning
Jingxuan Wei, Nan Xu, Guiyong Chang, Yin Luo, Bihui Yu, Ruifeng Guo
Shenyang Institute of Computing Technology, Chinese Academy of Sciences · Beijing Wenge Technology Company, Limited · University of Chinese Academy of Sciences
cs.CV, cs.AI
Submitted: 2024-04-02
Updated: 2026-08-11
Journal ref: Pattern Recognition 172 (2026) 112348
DOI: 10.1016/j.patcog.2025.112348
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 43/100
The gist: The paper introduces mChartQA, a "groundbreaking framework tailored for advanced multimodal chart question-answering" designed to address significant challenges posed by "intricate color patterns,
Terminology
Summary
The paper introduces mChartQA, a groundbreaking framework tailored for advanced multimodal chart question-answering
designed to address significant challenges posed by intricate color patterns, structural complexities, and implicit numerical data in charts.
The framework innovatively merges sophisticated language processing capabilities of LLMs with a state-of-the-art table-to-text engine, facilitating effective processing and integration of complex visual and textual information.
The mChartQA architecture consists of four main components:
-
Vision Encoder (E v): Processes a chart image I to
produce visual features V.
-
Connector (C):
Employs a cross-attention mechanism to align visual features V with the text encoder,
which iscrucial for correlating visual elements with corresponding textual data.
-
Chart-to-Text Engine (T):
Converts the chart image I into a textual representation T',
extracting key textual elements. -
Large Language Model (L):
Processes the tokenized question Q t, enhanced visual features V', and tokenized textual representation T't
to predict the answer A.
The training process is conducted in two stages:
-
Stage 1 - Visual-Language Alignment: This stage
focuses on training the Connector to optimize the alignment of visual and textual representations
using Captioning (utilizing datasets like COCO Caption, SBU, NoCaps, CC3M, and ShareGPT4V), Grounding (utilizing GRIT, Visual Genome, RefCOCO, RefCOCO+, and RefCOCOg), and Chart-to-text (utilizing ChartQA) tasks. -
Stage 2 - Visual-Language Reasoning: In this stage,
both the Connector and the Large Language Model are trained to enhance reasoning capabilities.
The model was evaluated on the public test sets of three datasets: ChartQA, PlotQA, and FigureQA. The evaluation focused on three main problem types: Color (requiring understanding of color theory
), Structure (relating to chart layout and structure
), and Textless (involving interpreting graphs containing implicit numerical data, where the graphical elements lack precise numerical values
).
Experimental results demonstrate that mChartQA showcases superior performance in tackling complex, multimodal chart question-answering tasks, especially in scenarios that have posed challenges for existing methods.
Specifically, the mChartQAIntern-LM2 version
demonstrated remarkable accuracy
and a remarkable ability to outperform existing models across these challenging scenarios,
particularly in the textless category across different datasets.
Ablation studies provided further insights:
-
Deplot Integration: The study found that
The absence of Deplot in either phase... leads to a noticeable decline in performance,
confirming itscrucial role in the model’s learning process.
-
Connector Type: Replacing the cross-attention connector with an
MLP-based approach (-Qformer + MLP)
resulted in lower performance inmost tasks, especially in complex reasoning scenarios.
-
Visual Encoder: The
ViT-448 encoder generally achieves superior results, particularly in handling complex chart structures and textless scenarios
compared to the ViT-384 encoder.
Error analysis identified specific patterns of failure:
-
Structure-Related Errors: These included
rounding off the answer,
language model hallucination, leading to a misspelling,
andinadequate structural recognition.
-
Color-Related Errors: These involved mistakes in
multi-step calculations,
limited color recognition accuracy,
and aneed for precise question interpretation.
-
Textless Chart Errors: These included
slight deviations in numerical estimations, precision mismatches between the model’s output and the standard answer, and confusion in the model’s understanding of the question scope.
Improvements for AI systems
1. Symbolic Reasoning and Formal Verification Layer
-
Improvement: Integrate a symbolic math engine and a post-processing verification module to intercept the LLM's numerical outputs.
-
Capability: The system will eliminate rounding errors, prevent numerical hallucinations (misspellings of numbers), and ensure that multi-step mathematical calculations are mathematically sound rather than just linguistically plausible.
2. Color-Semantic Attention Refinement
-
Improvement: Augment the Connector with a specialized color-attribute branch that utilizes high-resolution color histograms and color-space embeddings (e.g., LAB color space).
-
Capability: The system will achieve precise color recognition, allowing it to accurately map specific hues to legend entries and perform complex reasoning tasks that depend on distinguishing between similar color shades.
3. Coordinate-to-Value Regression Module
-
Improvement: Incorporate a regression head within the Chart-to-Text engine specifically designed to map pixel-level coordinates to axis-scale values through linear interpolation.
-
Capability: The system will minimize estimation deviations in textless charts by calculating precise numerical values based on visual position relative to the axes, reducing precision mismatches in the final output.
4. Hierarchical Structural Parsing Module
-
Improvement: Implement a multi-scale structural parser within the Vision Encoder that explicitly identifies chart components (axes, gridlines, tick marks, and data series boundaries) before passing features to the Connector.
-
Capability: The system will resolve structural recognition errors, enabling it to correctly interpret complex layouts such as multi-axis charts, grouped bar charts, and overlapping data series.
5. Intent-Driven Question Decomposition Module
-
Improvement: Add a pre-processing stage that uses a dedicated reasoning agent to decompose complex questions into a sequence of atomic sub-questions and logical constraints.
-
Capability: The system will prevent
question scope
confusion by ensuring every logical constraint in a query is addressed and by guiding the LLM through a structured step-by-step reasoning path.
Sources
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
- ShareGPT4V: Improving Large Multi-Modal Models with Better Captions
- PaLI-X: On Scaling up a Multilingual Vision and Language Model
- Microsoft COCO Captions: Data Collection and Evaluation Server
- PaLI-3 Vision Language Models: Smaller, Faster, Stronger
- Multimodal Document Analytics for Banking Process Automation
- FigureQA: An Annotated Figure Dataset for Visual Reasoning
- UniChart: A Universal Vision-language Pretrained Model for Chart Comprehension and Reasoning
- ChartAssisstant: A Universal Chart Multimodal Language Model via Chart-to-Table Pre-training and Multitask Instruction Tuning
- Kosmos-2: Grounding Multimodal Large Language Models to the World
- DOMINO: A Dual-System for Multi-step Visual Language Reasoning
- ChartX & ChartVLM: A Versatile Benchmark and Foundation Model for Complicated Chart Reasoning
- mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models