DEEPCHART: How Far are LLMs from Faithful Data-Science Chart Generation?

arXiv:2608.26757 · cs.AI · Submitted 2026-08-27 · Read on arXiv

cs.AI

Submitted: 2026-08-27

Updated: 2026-08-27

Code: https://github.com/tangdouer1005/DeepChart

Project page: https://moonshotai.github.io/Kimi-K2/thinking.html

License: http://creativecommons.org/licenses/by/4.0/

The gist: Faithful chart generation in real-world data-science workflows requires grounding visualizations in scattered evidence, computing chart-ready quantities, and rendering them accurately.

Terminology

Abstract

Faithful chart generation in real-world data-science workflows requires grounding visualizations in scattered evidence, computing chart-ready quantities, and rendering them accurately. Modern LLMs can produce visually plausible, instruction-compliant charts, yet data-level hallucinations remain difficult to detect in long, noisy, and multimodal contexts. To measure this gap, we introduce DEEPCHART, an expert-annotated benchmark of 1,482 task-conditioned chart-generation instances drawn from real-world scientific papers, financial filings, and ecosystem reports. DEEPCHART formulates chart generation as an Extract--Reason--Visualize pipeline and evaluates source-data extraction, derived-data reasoning, and chart rendering stage by stage. Experiments with state-of-the-art models show that visually plausible charts often conceal data-level hallucinations, with extraction and reasoning errors common in realistic long and multimodal settings. These findings suggest that larger context windows alone are insufficient; faithful chart generation also requires reliable evidence extraction and quantitative reasoning before rendering. Our benchmark and associated resources are available at https://github.com/tangdouer1005/DeepChart.

Sources

Related papers