VISTA: An End-to-End Benchmark for Visual Spec-to-Web-App Coding Agents

arXiv:2605.26144 · cs.SE, cs.AI, cs.CV · Submitted 2026-05-22 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "VISTA: An End-to-End Benchmark for Visual Spec-to-Web-App Coding Agents".

Jane: The gist The VISTA benchmark evaluates end-to-end web-app generation capabilities of LLM-based agents by targeting realistic UI-centric development from underspecified inputs.

Tom: First, who's behind it and why it matters.

Title and authors: Tom: So what’s this VISTA benchmark actually trying to do beyond just testing code generation? It’s about making sure these AI agents can handle the messiness of real web development.

Jane: Right, they are addressing a gap where most evaluations standardize the technology stack, which isn't how developers actually work in practice.

Lu: The paper sets up five specific scenarios to test this, varying both how visually accurate the input is and whether or not the agent has to stick to a certain technology stack.

Meng: It sounds like they are trying to see if an agent can make smart architectural decisions about which framework to use based on the task.

Lalam: They define these five conditions: you can have just text, or text with reference screenshots under specific stacks, or even screenshots plus a pruned Figma structure under one stack.

The paper's summary: Tom: The core of VISTA is that the agent has to do more than just write code; it has to interpret requirements, inspect the context, build visible pages with interactive behavior, and then actually run the application and fix things when they break.

Jane: It moves beyond static design-to-code by requiring the agent to handle running a multi-page app and making sure the visual intent is preserved throughout.

Lu: They are looking at how models interact with a software environment, not just their ability to output syntactically correct code or the next token.

Meng: So it’s about testing if an AI can connect front-end behavior to a data state, which is crucial for any real application development workflow.

Lalam: They are measuring structural alignment and behavioral completeness by looking at both how well the generated interface matches the reference and whether the controls actually work as expected.

The paper's improvements: Tom: The authors suggest a few ways to make this benchmark better, focusing on how we evaluate these agents. They propose using DOM-grounded reference matching along with behavior-specific browser tests.

Jane: That’s about checking if the generated controls actually exist as real elements in the code and if they behave right when you click them, which is a big step up from just looking at the final output.

Lu: They also argue that visual fidelity shouldn't be judged by a single metric; combining CLIP-based visual similarity with DOM grounding gives a more complete picture of structural alignment.

Meng: That means we need to ensure the agent isn't just making things *look* right from a distance, but that the underlying structure is solid and functional too.

Lalam: And they look at how agents edit their work, using something called Surgical Diff Score to see if they are doing localized patches instead of rewriting entire files.

Conclusion: Tom: So what’s the big picture here? VISTA is showing us that testing agents on realistic UI development tasks requires looking beyond just code quality and focusing on the entire interaction with a software environment.

Jane: It highlights the need for evaluations that measure structural alignment, functional behavior, and visual fidelity all at once across different input conditions.

Lu: The results show that performance isn't totally consistent across all these conditions; some setups actually make a bigger difference in how well agents perform than others.

Meng: From a practical standpoint, this suggests we need more robust ways to test if an agent can handle the real-world complexity of choosing the right framework and connecting different parts of an application.

Lalam: Ultimately, VISTA is giving us a clearer picture of what it means for AI agents to actually build something useful from a visual specification.

University of Arizona

cs.SE, cs.AI, cs.CV

Submitted: 2026-05-22

Updated: 2026-09-29

Comments: Project page: https://kaboider.github.io/VIS_APP/

Project page: https://www.figma.com/design

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 87/100

The gist: The gist The VISTA benchmark evaluates end-to-end web-app generation capabilities of LLM-based agents by targeting realistic UI-centric development from underspecified inputs.

Key concepts

Prompt-Information Conditions
These are five different ways to give an agent a task. They vary by how much visual detail is provided (from text only to screenshots with detailed Figma structures) and which programming technology stack the agent must use.
DOM-grounded Evaluator
This tool checks if the generated code matches the reference structure of a webpage. It verifies both that the elements are in the right place on the screen (structural alignment) and that they actually perform their intended function (behavioral completeness).
Structure-Function Score (S)
This metric measures task success based on critical interactions. It calculates a score by summing up how well an agent's interface structure aligns with the expected functions of those interactions, providing a single measure of overall correctness.
Surgical Diff Score
This quantifies the editing style of the agent. It measures whether an agent makes small, localized fixes (patches) or performs large, full-file rewrites to complete a task.

Terminology

Summary

The gist The VISTA benchmark evaluates end-to-end web-app generation capabilities of LLM-based agents by targeting realistic UI-centric development from underspecified inputs.

Benchmark Design and Conditions

VISTA defines five prompt-information conditions that vary along two axes: visual/structural fidelity and stack constraint The five prompt-information conditions that vary along two axes, visual/structural fidelity and stack constraint are (1) text only with free stack choice, (2) text with reference screenshots under three specified stacks, (3) text with reference screenshots under free stack choice, (4) text with screenshots and pruned Figma structure under a single specified stack, and (5) text with screenshots and pruned Figma structure under free stack choice. Table 1 shows that visual fidelity increases from C0 to C3/C4 (Figma structure), while Stack constraint varies independently.

Evaluation Methodology

To enable robust evaluation, each page in the benchmark is manually annotated with interactive UI components and around three visual anchor points addressing the well-known limitations of script-based testing tools such as Playwright in openended code generation settings. Evaluation combines DOM-grounded reference matching, behavior-specific browser tests, and CLIP-based visual similarity jointly measuring structural alignment, behavioral completeness, and overall visual fidelity. The DOM-grounded evaluator jointly measures whether the generated interface preserves the reference structure and whether the matched elements implement the expected behavior.

Evaluation Metrics

The evaluation utilizes several metrics to assess agent performance. The structure-function score S is defined as S = 1/N Sum X N i=1 Li · Bi, where N is the number of critical annotated interactions. Visual fidelity is assessed using CLIP image similarity as a complementary page-level metric rather than the sole evidence of design fidelity. Furthermore, agent trajectory and efficiency are analyzed through command usage, file edits, tool calls, skill usage, build attempts, test attempts, and failure-repair cycles. The Surgical Diff Score quantifies how much of an agent’s editing work is delivered through localized patches rather than full-file rewrites.

Key Findings on Fidelity and Correctness

The results indicate that visual fidelity and functional correctness are partially decoupled across both input conditions and agents. Adding screenshots (C0 → C1) raises mean localization from 0.489 to 0.577 (+18% relative) but slightly drops Combined (0.242 → 0.229). Removing the stack constraint while keeping the same visual inputs (C1 → C2) raises Combined from 0.229 to 0.264, the largest single jump in the table. The two strongest conditions are therefore the free-stack ones, C2 (best mean Combined) and C4 (best Combined median and CLIP, with Combined within 0.003 of C2).

Agent Workflow Analysis

The analysis of agent workflows suggests a shared macro-pattern where agents spend more of the early trajectory inspecting context, concentrate writing in the middle of the task, and increase verification toward the end. Claude-family models show stronger inspect→inspect transitions and substantially higher failure→inspect probabilities consistent with more repeated context gathering and a clearer return-to-diagnosis pattern after errors. Conversely, GPT-family models show more dispersed transitions after failure and verification suggesting a more heterogeneous repair loop or more frequent movement through auxiliary workflow states. The Surgical Diff Score has a weak negative Pearson correlation with the Combined structure-function score (ρ = −0.145).

Editing Style and Task Quality

The relationship between editing style and task success is largely orthogonal across models. Opus has the highest Combined (0.261) but the lowest Surgical Diff Score (22.4) and a rewrite share of 0.988 indicating that its strong outputs are produced almost entirely through rewrites. GPT-5.5 shows the opposite profile with the lowest Combined (0.223) but highest Surgical (33.6) and Strict (9.0) Diff Scores. This suggests that the model-level inversion reflects different editing strategies aligned with harness versions rather than a general dependence between surgical editing and task quality.

Limitations

The benchmark covers ten application categories, 128 pages, 3,253 interactive annotations, and four agent systems drawn from two harnesses. The three tech stacks evaluated under C1 and the single stack reused in C3 are also proposed by an LLM rather than a human panel which biases the benchmark toward stacks that LLMs find plausible. The DOM-grounded evaluator relies on a per-axis affine alignment, IoU and center-distance tiers, and interaction-specific behavior probes so we treat S as a robust ranking signal rather than an absolute measure of correctness. CLIP captures high-level visual resemblance but does not verify exact spacing, typography, color tokens, or accessibility properties. The Surgical Diff Score is computed from harnesslevel file-mutation events making it an agent-system-level property rather than a claim about model-internal preferences. All four values are well below the threshold at which one would expect editing style to predict task success.

References

The references for this paper are listed in the final section of the document. The references include work on program synthesis with large language models, agent-computer interfaces, and multimodal agents. The paper also cites foundational work on visual grounding and design-to-code benchmarks. The references are listed in the final section of the document. The references include work on program synthesis with large language models, agent-computer interfaces, and multimodal agents. The paper also cites foundational work on visual grounding and design-to-code benchmarks. The references are listed in the final section of the document. The references include work on program synthesis with large language models, agent-computer interfaces, and multimodal agents. The paper also cites foundational work on visual grounding and design-to-code benchmarks. The references are listed in the final section of the document. The references include work on program synthesis with large language models, agent-computer interfaces, and multimodal agents. The paper also cites foundational work on visual grounding and design-to-code benchmarks. The references are listed in the final section of the document. The references include work on program synthesis with large language models, agent-computer interfaces, and multimodal agents. The paper also cites foundational work on visual grounding and design-to-code benchmarks. The references are listed in the final section of the document. The references include work on program synthesis with large language models, agent-computer interfaces, and multimodal agents. The paper also cites foundational work on visual grounding and design-to-code benchmarks. The references are listed in the final section of the document. The references include work on program synthesis with large language models, agent-computer interfaces, and multimodal agents.

Improvements for AI systems

  1. Improve agent decision-making by incorporating explicit state management tools for task tracking, as demonstrated by Claude Code's stateful TodoWrite tool. This allows agents to maintain a structured checklist of progress (Explore inputs and choose stack, Scaffold Next.js app with TypeScript), leading to more deliberate and less exploratory workflows than those relying solely on natural-language agent message events.

  2. Enhance the ability of agents to produce high-quality, executable front-end code by systematically optimizing for structural alignment alongside functional correctness. This is achieved by jointly measuring DOM-grounded reference matching (measuring if controls exist as real DOM elements and appear near the corresponding reference locations) and behavior-specific browser tests, resulting in a higher structure-function score S.

  3. Improve visual fidelity in generated applications by utilizing CLIP-based visual similarity as a complementary metric to DOM grounding, ensuring agents prioritize overall visual resemblance while simultaneously verifying that annotated reference controls are present, grounded in the DOM, and structurally aligned with the mockup.

  4. Refine agent editing style by training models on minimizing rewrite-heavy workflows and maximizing localized modifications. This is achieved by using the Surgical Diff Score, which quantifies how much of an agent’s editing work is delivered through localized patches rather than full-file rewrites, aiming to reduce high rewrite shares observed in top-performing models.

  5. Improve stack selection accuracy by training agents to choose appropriate technologies based on application type and expected interaction patterns, as the benchmark uses LLMs to propose three suitable tech stacks per task. This moves agents beyond simply following a fixed stack template (C3) to making context-aware, task-dependent framework decisions.

Abstract

Coding agents can now build applications from design mockups, but a screen that looks right is not an application that works. Existing benchmarks either score visual reconstruction without a backend or test full-stack functionality from textual requirements, and few tie scores to the individual components a design requires. We introduce VISTA (VIsual Spec-To-App), an end-to-end benchmark in which coding agents turn multi-page design handoffs (Figma renders, structure, and textual requirements) into runnable full-stack Web and Mobile (Android) applications. VISTA provides 8,487 human-annotated interactive components across 18 applications (126 Web pages and 576 Android screens) and an executable evaluator that locates annotated components in the running application and probes their interactions, yielding localization, behavior, and joint scores traceable to individual design requirements. Across 14 deployed coding-agent systems, the best joint score is 0.553 on Web and 0.378 on Mobile, and 13 of 14 systems score lower on behavior than on localization on Web. Component probes detect unresponsive controls in 90.1% of audited Web deliveries. Development traces show where the process departs from the task: 44.5% of trajectories never open the design screenshots, 73.0% omit the prescribed self-audit, and 8.3% report success after a failed check. VISTA makes the gap between rendering a design and delivering a working application measurable at the level of individual components. Code is available at.

Sources

Related papers