From Visual Widgets to UI Code: Efficient Tool-Grounded Generation
Houston H. Zhang, Tao Zhang, Li Gu, Linfeng Ye, Yuanhao Yu, Xinxin Zuo, Yang Wang, Zhixiang Chi
McMaster University · University of Toronto · Concordia University
cs.CV, cs.LG
Submitted: 2026-08-12
Updated: 2026-08-14
Comments: ECCV2026 (MUCG Workshop)
Project page: https://zheny2751-dotcom.github.io/ui2code-n.github.io
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 95/100
The gist: Existing screenshot-to-code systems face a trade-off between flexibility and controllability.
Terminology
Summary
Existing screenshot-to-code systems face a trade-off between flexibility and controllability. Direct multimodal generation can hallucinate visible details, whereas structured pipelines reduce such errors through component-wise decomposition, predefined templates, and customized intermediate representations. These structures, however, introduce additional generative orchestration and restrict outputs to designs covered by the representation. The paper investigates whether selective tool grounding can improve the fidelity–efficiency trade-off of direct widget-to-code generation. The authors introduce WidgetGen, a lightweight tool-grounded framework that extracts observable text and color evidence, performs high-level layout and optional chart reasoning, and directly generates executable JavaScript XML (JSX). This design reduces reliance on component-wise generation while avoiding a fixed UI schema. Across six multimodal models and 1,000 held-out widgets, WidgetGen outperforms direct prompting and the structured Widget2Code pipeline on most visual reconstruction metrics, with consistent gains in area, legibility, and style. Finally, reconstruction-derived image-code pairs improve six Qwen-family open-weight models across every reported metric through supervised fine-tuning. These results establish WidgetGen as a strong lightweight baseline and show that selective evidence grounding offers an effective alternative to extensive representation constraints.
The key observation is that not every reconstruction subproblem requires generation. Visible text and dominant colors can be extracted from the screenshot, widget dimensions are known, and renderability can be checked through execution. Providing this evidence explicitly allows the MLLM to focus its capacity on the ambiguous decisions: interpreting layout, reconstructing charts, and synthesizing code. The authors hypothesize that selective evidence grounding can recover much of the controllability of a structured pipeline while reducing task-specific representation machinery. Based on this principle, WidgetGen extracts text and color evidence through optical character recognition (OCR) and palette analysis, performs high-level layout and optional chart reasoning, directly generates executable JSX, and renders the result in a browser. Unlike Widget2Code, it does not require explicit component-wise generation, a predefined template library, a customized intermediate UI representation, or a task-specific compiler. The framework therefore shifts control from constraining the output representation toward grounding the generator with observable input evidence, while retaining the flexibility of general-purpose UI code.
The contributions are threefold. First, the authors formulate selective evidence grounding as a design principle that separates observable visual facts from ambiguous generative decisions. Second, they instantiate this principle in WidgetGen and establish a lightweight alternative to component-wise, schema-constrained widget generation. Third, they provide a systematic evaluation across six MLLMs on 1,000 held-out widgets against both single-shot prompting and the structured Widget2Code pipeline. WidgetGen obtains the strongest results on most reconstruction metrics, with consistent improvements in Area, legibility, and style. As a secondary result, they show that reconstruction-derived image–code pairs provide effective supervision for fine-tuning open-weight Qwen models.
WidgetGen separates directly observable visual facts from decisions that require generative reasoning. Text and color evidence are first extracted from the screenshot. An MLLM then reasons about the whole-widget layout, optionally reconstructs chart content, and directly synthesizes executable React JSX. The resulting program is rendered once at the target dimensions. This design places explicit control on the evidence supplied to the generator rather than on a task-specific output representation. WidgetGen does not perform component-wise generation and does not require a predefined component inventory, a customized UI schema, or a task-specific compiler. The extracted evidence is shared across layout reasoning, chart reconstruction, and code synthesis, preserving a compact path from the screenshot to executable output.
The evidence layer targets visual facts that can be measured before code generation. For a screenshot I, OCR and palette analysis produce text evidence T = (t j, b j) where t j and b j are a recognized text string and its bounding box, and color evidence C = (c k, r k) where c k and r k are a hexadecimal color value and its approximate pixel coverage, with n C ≤ 8. The framework retains the observed dimensions (W, H) for terminal rendering. The text extractor localizes visible strings and represents each observation with its recognized content and bounding box, supporting both lexical fidelity and spatial placement. The palette extractor summarizes the screenshot using dominant colors and their approximate pixel coverage, capturing the principal background, foreground, and accent colors while suppressing minor pixel-level variation.
WidgetGen conditions each generative decision on the original screenshot and the extracted evidence. The layout reasoner produces a high-level description M L = f L(I, T, C) that identifies the main spatial regions and whether the widget contains a chart. This reasoning operates on the widget as a whole rather than predicting a sequence of components from a predefined vocabulary. When M L indicates a chart, a dedicated call generates chart JSX; otherwise the call is skipped. When present, the resulting chart code is supplied as context to the final generator. Separating this call focuses generative capacity on chart geometry only when needed. The final generator jointly receives the screenshot, extracted evidence, layout description, and optional chart code, producing executable JSX directly, without translating through component templates or a customized intermediate UI language. The browser renderer executes the program at the observed dimensions, verifying that the generated artifact can be executed and producing the image used for evaluation.
Each successfully executed reconstruction yields an aligned image–code pair for supervised fine-tuning. Given a generated program, browser execution produces a rendered image; the rendered image, its dimensions, and the OCR and palette evidence recomputed from it form the model input, while the program is the target JSX. Because the target program is executable by construction, these pairs require no manual code annotation and have an unambiguous correspondence between image and code. Moreover, since the programs originate from reconstructing visual designs rather than unconstrained code synthesis, the resulting pairs retain layout and appearance patterns from the target task. These pairs are used to fine-tune open-weight MLLMs for evidence-conditioned JSX generation.
The evaluation uses the widget2code-benchmark introduced by Widget2Code, a collection of widget screenshots split into a train set of 1,822 widgets and a test set of 1,000 widgets. Every generated widget is evaluated by rendering its JSX program back to a PNG and comparing the rendered image against the ground-truth screenshot using nine metrics adopted from Widget2Code: Content Aspect Ratio Similarity (Content), Area Ratio Similarity (Area), OCR Text Jaccard (Text), Local Contrast Similarity (LocCon), Palette Distance (Palette), Vibrancy Consistency (Vibrancy), SSIM, LPIPS, and Geometry. Layout, legibility, style, and geometry metrics are reported on [0,100] where larger is better; SSIM is larger-is-better on [0,1], and LPIPS is smaller-is-better. Three methods are compared using the same six multimodal models: Single-shot Prompting, Widget2Code, and WidgetGen. The models are GPT-4o, Gemini 3 Pro, Claude Opus 4.5, Seed 2.0 Pro, Qwen3-VL-Plus, and Qwen3.5-Plus. Across all six configurations, only the language model changes; the tools and renderer remain the same. Text evidence is extracted using EasyOCR on a single H200 GPU. For palette evidence, octree color quantization with 16 initial clusters is applied, visually similar colors are merged, colors covering less than 3% of the image are discarded, and up to eight colors ranked by coverage are retained. The renderer uses a four-page browser worker pool and executes each generated widget at its target dimensions.
The main results show three patterns. WidgetGen improves legibility and style over both baselines for every model, reaches a geometry score of 100.00 for every model, and obtains the best Area score for all six models. It also obtains the best Content score for four of six models, with Widget2Code higher only for Qwen3.5-Plus and Claude Opus 4.5. Geometry is nearly saturated for both structured methods and is therefore not the main source of separation. The metric pattern is consistent with their different control mechanisms: Widget2Code constrains generation through a DSL and compiler, whereas WidgetGen supplies measured text and color evidence before direct JSX synthesis. The Content exceptions indicate that representation-level structure remains useful for some backbones.
The fidelity–efficiency trade-off is profiled with the same backbone and benchmark, recording inference usage, specialized-tool overhead, end-to-end latency, and rendering reliability. To distinguish evidence grounding from simply using more model calls, the comparison includes a call-matched multi-stage baseline that performs layout and chart reasoning without OCR or palette evidence. Palette extraction and OCR run in parallel; the reported WidgetGen tool time is the maximum wall-clock time of the two tools, while the single-shot and no-tools variants do not run either tool. API cost is computed using the official model pricing and reported in U.S. dollars. Overall, WidgetGen provides a favorable fidelity–efficiency trade-off, improving reconstruction quality while reducing inference cost relative to the structured pipeline. For Qwen3-VL-Plus, WidgetGen uses 2.27 MLLM calls per widget, 14,909 input tokens, 2,607 output tokens, 17,516 total tokens, 3.88 seconds specialized-tool time, 98.9/271.3 seconds end-to-end latency p50/p95, 5.87 USD API cost per 1,000 widgets, 98.0% render success, 0.71 SSIM, and 0.34 LPIPS. In contrast, Widget2Code uses 3.47 calls, 39,982 input tokens, 2,035 output tokens, 42,017 total tokens, 24.5 seconds tool time, 112/363 seconds latency, 8.38 USD cost, 97.0% render success, 0.70 SSIM, and 0.35 LPIPS. Single-shot prompting uses 1.00 call, 807 input tokens, 1,202 output tokens, 2,009 total tokens, 0 tool time, 31.5/81.3 seconds latency, 1.84 USD cost, 86.0% render success, 0.61 SSIM, and 0.44 LPIPS.
The evidence ablation isolates the effects of extracted evidence and high-level reasoning using the same model and evaluation setup. The call-matched no-tools variant retains layout and optional chart reasoning but removes OCR and palette evidence; the remaining variants add one evidence source at a time or remove chart reasoning from the full system. Chart reasoning is additionally evaluated on the chart-containing subset so that its effect is not diluted by widgets for which the call is skipped. Compared with single-shot prompting, the multi-stage variant without tools improves the reconstruction metrics, indicating that high-level reasoning contributes to reconstruction quality. Adding text and color evidence produces further gains in legibility and style, with the full configuration achieving the best Text, Palette, and Geometry scores. On the chart subset, chart reasoning improves several content and appearance metrics. For the All subset with Qwen3-VL-Plus, single-shot achieves Content 16.65, Area 66.22, Text 58.32, LocCon 58.59, Palette 44.30, Vibrancy 39.49, SSIM 0.610, LPIPS 0.440, Geometry 84.65; multi-stage no tools achieves Content 26.58, Area 85.42, Text 70.05, LocCon 69.44, Palette 61.43, Vibrancy 58.40, SSIM 0.690, LPIPS 0.350, Geometry 99.99; OCR only achieves Content 25.71, Area 85.71, Text 71.58, LocCon 69.62, Palette 61.31, Vibrancy 58.42, SSIM 0.707, LPIPS 0.340, Geometry 99.99; Palette only achieves Content 27.18, Area 85.78, Text 70.00, LocCon 70.18, Palette 61.44, Vibrancy 59.05, SSIM 0.709, LPIPS 0.339, Geometry 99.99; Full WidgetGen achieves Content 27.60, Area 85.94, Text 72.20, LocCon 70.55, Palette 61.62, Vibrancy 58.86, SSIM 0.710, LPIPS 0.338, Geometry 100.00. For the Charts subset, w/o chart reasoning achieves Content 25.59, Area 87.20, Text 71.09, LocCon 72.74, Palette 60.21, Vibrancy 54.92, SSIM 0.723, LPIPS 0.339, Geometry 99.95; Full WidgetGen achieves Content 27.49, Area 86.52, Text 72.53, LocCon 72.51, Palette 60.24, Vibrancy 56.66, SSIM 0.725, LPIPS 0.336, Geometry 100.00.
Qualitative analysis compares the three methods on a chart, an icon-dense grid, and a compact text widget. In the chart example, single-shot prompting reconstructs only part of the plot and Widget2Code omits most chart content, whereas WidgetGen recovers the principal curves, axes, legend, and annotation while retaining visible shape differences from the target. The icon-grid example shows improved preservation of the global composition and palette, but also exposes residual errors in icon identity. In the compact time widget, WidgetGen more closely matches the target’s typographic weight and text placement while preserving the overall composition.
As a secondary study, the authors evaluate whether the executable reconstructions provide useful task-specific supervision for open-weight vision-language models. This experiment is separate from the main system comparison: it studies WidgetGen as a data-generation process rather than as an inference-time replacement for the base model. WidgetGen with Gemini 3.1 Pro is applied to the 1,822 training widgets. Each executable program is rendered to produce its paired image, from which text and color evidence are re-extracted. Pairing a program with its own rendering preserves image–code correspondence even when the reconstruction differs from the source screenshot. After quality filtering and regeneration, 1,822 executable image–code pairs are obtained, one for each training widget. Six checkpoints from the Qwen3-VL, Qwen3.5, and Qwen3.6 families, ranging from 4B to 32B parameters, are evaluated. Each model is adapted with LoRA while keeping the vision encoder frozen. Rank-32 adapters are trained for four epochs with a scaling factor of 64, dropout 0.05, learning rate 10-4, and effective batch size 16. Base and adapted checkpoints use the same evidence-conditioned inputs and decoding setup.
Table 4 shows consistent improvements across all six checkpoints. Reconstruction-derived supervision improves every reported metric, with gains spanning layout, legibility, style, and perceptual similarity. Within this evaluation, the adapted 27B checkpoints also exceed the proprietary single-shot baselines on most metrics. The process is similar to distilling the knowledge from stronger Gemini to weaker models. For example, Qwen3-VL-4B-Instruct improves from Content 19.09 to 26.01, Area 53.24 to 80.03, Text 51.26 to 66.01, LocCon 41.49 to 70.10, Palette 43.51 to 60.28, Vibrancy 39.57 to 56.70, SSIM 0.64 to 0.70, LPIPS 0.36 to 0.34, Geometry 60.93 to 100.00. Qwen3-VL-8B-Instruct improves from Content 22.72 to 29.97, Area 58.05 to 80.99, Text 58.12 to 68.80, LocCon 43.84 to 70.43, Palette 44.90 to 59.73, Vibrancy 42.47 to 57.83, SSIM 0.70 to 0.71, LPIPS 0.34 to 0.33, Geometry 65.79 to 100.00. Qwen3-VL-32B-Instruct improves from Content 17.03 to 34.85, Area 57.34 to 84.86, Text 66.14 to 73.51, LocCon 58.34 to 73.17, Palette 47.21 to 62.95, Vibrancy 43.10 to 61.11, SSIM 0.70 to 0.72, LPIPS 0.34 to 0.31, Geometry 95.24 to 99.71. Qwen3.5-9B improves from Content 20.15 to 36.85, Area 64.34 to 86.33, Text 55.60 to 72.89, LocCon 44.57 to 73.19, Palette 46.58 to 64.37, Vibrancy 42.80 to 63.82, SSIM 0.69 to 0.73, LPIPS 0.35 to 0.29, Geometry 55.57 to 100.00. Qwen3.5-27B improves from Content 27.68 to 41.92, Area 80.65 to 87.91, Text 63.87 to 75.89, LocCon 51.13 to 76.75, Palette 50.39 to 65.90, Vibrancy 47.32 to 66.36, SSIM 0.67 to 0.74, LPIPS 0.35 to 0.27, Geometry 55.95 to 100.00. Qwen3.6-27B improves from Content 30.53 to 42.80, Area 82.22 to 88.84, Text 65.41 to 75.67, LocCon 49.54 to 76.20, Palette 52.01 to 67.30, Vibrancy 47.47 to 68.04, SSIM 0.69 to 0.74, LPIPS 0.34 to 0.26, Geometry 50.35 to 100.00.
A supervision-route comparison on Qwen3-VL-32B-Instruct using 1,822 training pairs per route and the same nine metrics shows that WidgetGen supervision achieves the best result on eight of the nine metrics. None (base checkpoint) achieves Content 17.03, Area 57.34, Text 66.14, LocCon 58.34, Palette 47.21, Vibrancy 43.10, SSIM 0.70, LPIPS 0.34, Geometry 95.24. WidgetFactory-generated pairs achieve Content 24.15, Area 75.81, Text 68.95, LocCon 62.43, Palette 52.61, Vibrancy 46.13, SSIM 0.709, LPIPS 0.343, Geometry 100.00. WidgetGen reconstruction pairs achieve Content 34.85, Area 84.86, Text 73.51, LocCon 73.17, Palette 62.95, Vibrancy 61.11, SSIM 0.72, LPIPS 0.31, Geometry 99.71. The comparison indicates that reconstruction-derived JSX pairs provide more effective supervision for the evaluated direct-generation setting than the size-matched DSL-based data.
The conclusion states that WidgetGen, a lightweight tool-grounded framework for visual widget-to-code generation, combines extracted text and color evidence with high-level layout and optional chart reasoning before directly synthesizing executable JSX. Across six multimodal models and 1,000 held-out widgets, WidgetGen improves most reconstruction metrics over both direct prompting and Widget2Code while providing a favorable fidelity–efficiency trade-off. Ablations show complementary contributions from high-level reasoning and visual evidence. In addition, reconstruction-derived image–code pairs improve six open-weight Qwen-family models across all reported metrics. These results establish WidgetGen as a strong and efficient baseline for widget-to-code generation and demonstrate the value of grounding direct code synthesis with task-relevant visual evidence. The limitations are that the current evaluation focuses on static widgets from a single dataset, so generalization to broader UI distributions remains to be established; the metrics assess rendered visual fidelity rather than interactive behavior, accessibility, or code maintainability; and the generated-supervision study is limited to Qwen-family models, with extending it to other model families remaining future work.
Improvements for AI systems
Based on this paper, I can improve AI systems in the following ways:
-
Implement selective evidence grounding in multimodal generation systems: Instead of forcing models to generate all details from scratch, extract observable facts (e.g., text via OCR, dominant colors via palette analysis, dimensions) and feed them as explicit inputs. This reduces hallucination and improves fidelity, as shown by WidgetGen’s gains in area, legibility, and style metrics across six models.
-
Add a layout-and-chart reasoning pre-step before code generation: Before synthesizing final output, have the model produce a high-level spatial description and detect whether a chart is present. This improves reconstruction quality (e.g., Content score from 16.65 to 27.60 in the ablation) and can be conditionally skipped when not needed, saving compute.
-
Use executable renderings as self-supervised training data: Generate image–code pairs by rendering the system’s own outputs, then fine-tune smaller models on these pairs. This improved all six Qwen-family models on every metric (e.g., Qwen3.5-27B LPIPS from 0.35 to 0.27, Geometry from 55.95 to 100.00), providing a scalable data-generation method without manual annotation.
-
Optimize token efficiency by separating evidence extraction from generation: By providing OCR and palette data as structured inputs, the system reduces input tokens (e.g., from 39,982 in Widget2Code to 14,909 in WidgetGen for Qwen3-VL-Plus) and lowers API cost by 30% while improving quality—useful for cost-sensitive deployment.
-
Improve render success rates through tool-grounded verification: Use a browser renderer to check executability, raising render success from 86.0% (single-shot) to 98.0% (WidgetGen). This can be applied to any code-generation system to filter invalid outputs before user-facing use.
-
Enhance chart-specific generation with dedicated sub-calls: When the layout reasoner detects a chart, route to a specialized generation step; this improved chart-related metrics (e.g., Vibrancy from 54.92 to 56.66) without affecting non-chart widgets, allowing targeted capacity allocation.
-
Distill knowledge from stronger models via reconstruction: Use a powerful model (e.g., Gemini 3.1 Pro) to generate reconstructions, then fine-tune smaller models on those pairs. This lifted Qwen3-VL-32B’s Content from 17.03 to 34.85, surpassing proprietary single-shot baselines—enabling high-quality performance with open-weight models.
-
Reduce hallucination of visible details by grounding in measured evidence: The system’s design—where text and colors are extracted, not guessed—led to consistent improvements in Text and Palette metrics (e.g., Text from 58.32 to 72.20 in the full system). This principle can be applied to any vision-to-code or vision-to-text task where ground-truth visual properties are measurable.
What the improved AI system can do:
-
Generate UI code (JSX, HTML, etc.) from screenshots with higher visual fidelity (better area, legibility, style, and geometry) than both direct prompting and structured pipelines, while using fewer tokens and lower cost.
-
Automatically extract and use observable visual evidence (text, colors, dimensions) to reduce errors, and verify outputs by rendering them in a browser to ensure executability.
-
Self-improve by generating its own training data from reconstructions, enabling smaller models to reach or exceed proprietary model performance on widget-to-code tasks.
-
Handle charts and complex layouts more accurately by conditionally invoking specialized reasoning steps, without sacrificing efficiency on simpler inputs.
-
Serve as a lightweight, flexible baseline that avoids rigid templates or DSLs, making it adaptable to new UI designs without retraining the representation layer.
Abstract
Existing screenshot-to-code systems face a trade-off between flexibility and controllability. Direct multimodal generation can hallucinate visible details, whereas structured pipelines reduce such errors through component-wise decomposition, predefined templates, and customized intermediate representations. These structures, however, introduce additional generative orchestration and restrict outputs to designs covered by the representation. We investigate whether selective tool grounding can improve the fidelity--efficiency trade-off of direct widget-to-code generation. We introduce WidgetGen, a lightweight tool-grounded framework that extracts observable text and color evidence, performs high-level layout and optional chart reasoning, and directly generates executable JavaScript XML (JSX). This design reduces reliance on component-wise generation while avoiding a fixed UI schema. Across six multimodal models and 1, 000 held-out widgets, WidgetGen outperforms direct prompting and the structured Widget2Code pipeline on most visual reconstruction metrics, with consistent gains in area, legibility, and style. Finally, reconstruction-derived image-code pairs improve six Qwen-family open-weight models across every reported metric through supervised fine-tuning. These results establish WidgetGen as a strong lightweight baseline and show that selective evidence grounding offers an effective alternative to extensive representation constraints.
Sources
- Qwen3-VL Technical Report
- Advancing vision-language models in front-end development via data synthesis
- Distilling the Knowledge in a Neural Network
- Unlocking the conversion of Web Screenshots into HTML Code with the WebSight Dataset
- DesignBench: A Comprehensive Benchmark for MLLM-based Front-end Code Generation
- UI-UG: A Unified MLLM for UI Understanding and Generation
- ASMIL: Attention-Stabilized Multiple Instance Learning for Whole Slide Imaging
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models