Beyond Static Charts: Can Language and Vision Language Models Generate Interactive Data Visualization Interfaces?
cs.CL
Submitted: 2026-08-11
Updated: 2026-08-11
Code: https://github.com/vis-nlp/VIS-GEN
License: http://creativecommons.org/licenses/by/4.0/
The gist: Data visualization is central to analytical reasoning, but real-world analysis increasingly requires language-driven interactive interfaces rather than static charts.
Terminology
Abstract
Data visualization is central to analytical reasoning, but real-world analysis increasingly requires language-driven interactive interfaces rather than static charts. Although recent large language and vision language models (LLMs/VLMs) have shown promise in generating static charts from natural language, their ability to generate interactive data visualization interfaces remains largely unexplored due to the lack of benchmarks. We introduce VIS-GEN, a benchmark for evaluating how well LLMs/VLMs can generate interactive visualization interfaces from natural language queries. VIS-GEN comprises 3,042 samples covering diverse analytical intents, including data filtering, temporal analysis, and visualization editing, each paired with dataset metadata and natural language queries that are designed to reflect realistic, goal driven data exploration scenarios. We benchmark 14 state-of-the-art open-source and closed-source LLMs/VLMs, revealing large performance gaps and frequent failures on queries involving implicit intent, multiple interaction alternatives, and complex editing operations, highlighting interactive interface generation as a key open challenge beyond static chart synthesis. To address this, we propose a structured multi stage interface generation framework that decomposes the task into visualization design representation, generation of multiple interface candidates, constraint-aware critique, and self-refinement. This approach improves the best models pass rate by 15.9 percentage points, demonstrating a practical path toward more reliable language-driven interactive visualization systems. We release VIS-GEN at https://github.com/vis-nlp/VIS-GEN.
Sources
- Generative Interfaces for Language Models
- The Llama 3 Herd of Models
- ChartLlama: A Multimodal LLM for Chart Understanding and Generation
- Mistral 7B
- Unlocking the conversion of Web Screenshots into HTML Code with the WebSight Dataset
- nvBench: A Large-Scale Synthesized Dataset for Cross-Domain Natural Language to Visualization Task
- GPT-4 Technical Report
- LLM-Based Data Science Agents: A Survey of Capabilities, Challenges, and Future Directions
- Code Llama: Open Foundation Models for Code
- Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
- Seq2SQL: Generating Structured Queries from Natural Language using Reinforcement Learning
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering