DEEPCHART: How Far are LLMs from Faithful Data-Science Chart Generation?
cs.AI
Submitted: 2026-08-27
Updated: 2026-08-27
Code: https://github.com/tangdouer1005/DeepChart
Project page: https://moonshotai.github.io/Kimi-K2/thinking.html
License: http://creativecommons.org/licenses/by/4.0/
The gist: Faithful chart generation in real-world data-science workflows requires grounding visualizations in scattered evidence, computing chart-ready quantities, and rendering them accurately.
Terminology
Abstract
Faithful chart generation in real-world data-science workflows requires grounding visualizations in scattered evidence, computing chart-ready quantities, and rendering them accurately. Modern LLMs can produce visually plausible, instruction-compliant charts, yet data-level hallucinations remain difficult to detect in long, noisy, and multimodal contexts. To measure this gap, we introduce DEEPCHART, an expert-annotated benchmark of 1,482 task-conditioned chart-generation instances drawn from real-world scientific papers, financial filings, and ecosystem reports. DEEPCHART formulates chart generation as an Extract--Reason--Visualize pipeline and evaluates source-data extraction, derived-data reasoning, and chart rendering stage by stage. Experiments with state-of-the-art models show that visually plausible charts often conceal data-level hallucinations, with extraction and reasoning errors common in realistic long and multimodal settings. These findings suggest that larger context windows alone are insufficient; faithful chart generation also requires reliable evidence extraction and quantitative reasoning before rendering. Our benchmark and associated resources are available at https://github.com/tangdouer1005/DeepChart.
Sources
- Qwen3-VL Technical Report
- ChartAB: A Benchmark for Chart Grounding & Dense Alignment
- FinQA: A Dataset of Numerical Reasoning over Financial Data
- PaddleOCR-VL: Boosting Multilingual Document Parsing via a 0.9B Ultra-Compact Vision-Language Model
- DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
- LIDA: A Tool for Automatic Generation of Grammar-Agnostic Visualizations and Infographics using Large Language Models
- Data2Vis: Automatic Generation of Data Visualizations Using Sequence to Sequence Recurrent Neural Networks
- Benchmarking Retrieval-Augmented Multimodal Generation for Document Question Answering
- VisEval: A Benchmark for Data Visualization in the Era of Large Language Models
- Data Interpreter: An LLM Agent For Data Science
- VizML: A Machine Learning Approach to Visualization Recommendation
- InfiAgent-DABench: Evaluating Agents on Data Analysis Tasks
- DePlot: One-shot visual language reasoning by plot-to-table translation
- DA-Code: Agent Data Science Code Generation Benchmark for Large Language Models
- nvBench: A Large-Scale Synthesized Dataset for Cross-Domain Natural Language to Visualization Task
- Deep Research Agents: A Systematic Examination And Roadmap
- Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
- DSBench: How Far Are Data Science Agents from Becoming Data Science Experts?
- DVQA: Understanding Data Visualizations via Question Answering
- $C^2$: Scalable Auto-Feedback for LLM-based Chart Generation
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection