VisEditBench: Can Vision-Language Models Edit Visualization Code from Multimodal Feedback?

arXiv:2608.10408 · cs.CL · Submitted 2026-08-11 · Read on arXiv

Mizanur Rahman, Arshia Azimlu, Shadikur Rahman, Md Tahmid Rahman Laskar, Amran Bhuiyan, Shafiq Joty, Enamul Hoque Prince

York University · Nanyang Technological University · Salesforce AI Research

cs.CL

Submitted: 2026-08-11

Updated: 2026-08-12

Code: https://github.com/vis-nlp/VisEditBench

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: VisEditBench is a benchmark for evaluating vision-language models (VLMs) on the task of editing visualization code from multimodal feedback, addressing a gap in existing benchmarks that primarily

Terminology

Summary

VisEditBench is a benchmark for evaluating vision-language models (VLMs) on the task of editing visualization code from multimodal feedback, addressing a gap in existing benchmarks that primarily focus on generating visualizations from scratch. The paper introduces a dataset of 1,395 human-annotated visualization code-editing tasks, covering two practical settings: feedback-guided repair, where models revise code using buggy or human-marked charts with textual feedback, and reference-guided restyling, where models modify code to match a target chart image while preserving data semantics. The tasks are grounded in realistic workflows, collected from Stack Overflow, Matplotlib and Vega-Lite issue reports, and real-world datasets, and span eight editing intents: correctness repair, quality improvement, robustness/generalization, style adaptation, constraint satisfaction, consistency harmonization, refactor/transformation, and style-aware error repair. The dataset includes 1,156 Matplotlib and 239 Vega-Lite examples, with 57.1% of tasks being multi-cause, and difficulty levels ranging from easy (57.5%) to hard (27.3%).

The paper evaluates 20 state-of-the-art VLMs in a zero-shot setting using metrics of code executability, task accuracy, readability and clarity, visual quality, visual similarity, and a strict final pass rate. Results show that visualization code editing remains challenging: Claude-4.6-Sonnet achieves the best overall pass rate at 74.46%, followed by GPT-5 at 65.68% and Claude-4.5-Sonnet at 62.73%, while most open-source models remain below 50%, with Qwen3-VL-32B being the best open-source model at 51.72%. Performance is particularly weak on visually grounded style adaptation, where even Claude-4.6-Sonnet achieves only 55.71% and GPT-4o only 10.00%. The paper notes that executability alone is not the primary bottleneck, as strong models execute successfully on over 93% of examples but achieve lower visual similarity scores, indicating failures in preserving visual semantics and stylistic alignment.

To establish a strong baseline, the paper proposes VisEditAgent, a render-grounded editing framework that iteratively plans, generates multiple candidate code revisions, executes and renders them, validates the outputs visually, and refines the selected solution through feedback. Using GPT-4o as the base model, VisEditAgent improves the overall pass rate from 55.75% to 67.99%, with the largest gains on visually grounded tasks: style adaptation improves from 10.00% to 62.85%, consistency harmonization from 47.46% to 64.41%, and refactor/transformation from 42.86% to 60.32%. Ablation studies show that removing multi-candidate generation lowers the pass rate from 67.99% to 61.94%, and removing refinement lowers it to 59.86%, demonstrating the importance of both candidate selection and feedback-based refinement. Human evaluation on all 1,395 examples for GPT-4o and a 500-example stratified sample for Qwen3-VL-4B confirms the improvements, with Pearson correlations between automatic and human judgments ranging from 83.38 to 87.00 across scalar metrics.

Error analysis identifies four recurring failure modes: executable but visually incorrect outputs, weak grounding in visual feedback, poor reference-style matching, and incorrect or incomplete transformations. The paper concludes that reliable visualization editing requires iterative multimodal reasoning rather than single-pass code generation, and that VisEditBench and VisEditAgent establish a foundation for advancing visually grounded, feedback-aware visualization authoring systems. Limitations include coverage of only Matplotlib and Vega-Lite libraries, not Plotly, D3.js, or ggplot2, and the benchmark is not intended to exhaustively cover all visualization-editing needs.

Improvements for AI systems

Improvements to AI systems:

  1. Add a render-grounded iterative refinement loop. Instead of generating code once, the AI system should generate multiple candidate code revisions, execute and render them, validate the visual output against the input chart or feedback, and then refine the best candidate based on visual discrepancies. This turns single-pass code generation into a closed-loop, self-correcting process.

  2. Implement multi-candidate generation with visual selection. The system should produce several distinct code variants (e.g., 3–5) and rank them not just by executability but by visual similarity to the target or by how well they address the textual feedback. This improves robustness, especially for style adaptation and consistency tasks where a single attempt often fails.

  3. Add explicit visual-semantic validation metrics. The system should compare rendered outputs against the reference image using perceptual similarity (e.g., SSIM, CLIP-based embeddings) and structural checks (e.g., axis ranges, legend presence, color mapping) to catch cases where code runs but the chart is visually wrong. This addresses the failure mode of executable but visually incorrect.

  4. Incorporate feedback-driven error classification. The system should classify the type of failure (e.g., style mismatch, incorrect transformation, missing constraint) after rendering and use that classification to guide the next refinement step, rather than blindly regenerating code. This improves efficiency and success on multi-cause tasks.

  5. Enhance visual grounding in the model’s reasoning. The AI should be trained or prompted to explicitly describe what it sees in the input chart (e.g., the bars are blue, the y-axis is log-scaled, the legend is missing) before generating edits. This reduces weak grounding in visual feedback, a key failure mode.

  6. Add a style-transfer module for reference-guided restyling. The system should separate data semantics from visual styling (e.g., color palette, font, gridlines, marker shapes) and apply transformations only to the styling layer while preserving data mappings. This directly targets the poor performance on style adaptation.

  7. Implement a consistency harmonizer for multi-chart outputs. When editing one chart to match another, the system should extract a style profile from the reference chart (colors, fonts, axis formats) and apply it uniformly across all charts in the same figure or dashboard, improving consistency harmonization.

  8. Add a transformation verifier. For refactor/transformation tasks, the system should check that the new code produces the same underlying data values and relationships as the original, even if the visual representation changes. This prevents incorrect or incomplete transformations.

  9. Use difficulty-aware prompting and fallback strategies. For hard tasks (27.3% of the benchmark), the system should automatically trigger more candidate generation and more refinement iterations, while for easy tasks it can use a single pass to save compute. This improves overall pass rate without excessive cost.

  10. Integrate a human-feedback simulation module. The system should be able to take natural language feedback (e.g., make the bars thinner, the legend overlaps the plot) and convert it into specific code-level changes (e.g., adjust bar width parameter, reposition legend) using a structured mapping, improving feedback-guided repair.


What the improved AI system can do:

  • Self-correct visualization code by rendering, visually inspecting, and refining its own output until it matches the intended chart or feedback, achieving pass rates comparable to the best models (e.g., 68% on VisEditBench with GPT-4o base).

  • Handle visually grounded style edits (e.g., make it look like this reference chart) with high accuracy, improving from 10% to 63% on style adaptation tasks.

  • Maintain data integrity while changing visual appearance, ensuring that transformations do not alter underlying data semantics.

  • Work across multiple chart libraries (Matplotlib, Vega-Lite) and generalize to others like Plotly or D3.js with minimal adaptation.

  • Provide explainable failure analysis by identifying whether a failure is due to execution, visual mismatch, style mismatch, or transformation error, and then targeting the fix accordingly.

  • Operate in zero-shot mode without fine-tuning, making it deployable immediately for real-world visualization editing workflows, such as debugging charts from issue reports or restyling dashboards to match brand guidelines.

Abstract

Vision-language models (VLMs) have shown strong capabilities in generating visualization code from textual or visual specifications. However, real-world visualization authoring is inherently iterative: users frequently revise existing visualizations to repair flawed charts or adapt them to desired styles. Existing benchmarks primarily evaluate generation from scratch, leaving visualization code editing from multimodal feedback largely unexplored. We introduce VisEditBench, a benchmark of 1,395 human-annotated visualization code-editing tasks grounded in realistic visualization workflows and failure cases. VisEditBench covers two practical settings: feedback-guided repair, where models revise visualization code using buggy or marked charts together with textual feedback, and reference-guided restyling, where models modify code to match a target chart image. Evaluating 20 state-of-the-art VLMs reveals that visualization code editing remains challenging: Claude-4.6-Sonnet achieves the best overall pass rate of 74.46%, while most open-source models remain below 50%. Performance is particularly weak on visually grounded style adaptation, where Claude-4.6-Sonnet achieves only 55.71%. To establish a strong baseline, we further propose VisEditAgent, a render-grounded editing framework that iteratively generates, executes, validates, and refines candidate edits. Built on GPT-4o, VisEditAgent improves overall pass rate from 55.75% to 67.99%, demonstrating the importance of render-grounded feedback for faithful visualization editing. We will release VisEditBench at https://github.com/vis-nlp/VisEditBench.

Sources

Related papers