UI2App: Benchmarking Visual Interaction Inference in Executable Web Application Generation
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "UI2App: Benchmarking Visual Interaction Inference in Executable Web Application Generation".
Jane: Large language models have demonstrated growing competence in web page generation, but existing text-driven approaches are limited by complex prompts, and image-driven paradigms lack systematic evaluation of interaction capabilities.
Tom: First, who's behind it and why it matters.
Paper summary: Tom: So, to recap where we are with UI2App: this benchmark consists of three hundred twenty-seven screenshots grouped into forty-five sets that represent runnable multi-route web applications <ref:2607.06306#pg0,327 screenshots grouped into 45>. The authors set out to create a protocol that assesses artifacts across four dimensions: executability, navigation reachability, visual fidelity, and interaction inference.
Jane: It really tackles the gap where we have image-driven approaches that focus too much on how the app looks instead of whether it actually works when you interact with it.
Lu: The central thesis they are pushing is that inferring complete interaction behavior from static screenshots alone remains a significant challenge for models, especially because cross-route state persistence seems to be a major bottleneck across the frontier.
Meng: They introduce this framework to systematically evaluate these interaction capabilities using an "interaction taxonomy" of seven categories, which structures how interactions are assessed into UI-state, data-state, and cross-route persistence.
Lalam: That taxonomy approach is key because it allows them to do a structured assessment of inferred behaviors without needing a unique ground-truth execution trace, which is a huge practical advantage for testing.
Conclusion: Tom: The authors of UI2App are Grace Man Chen, Litao Guo, Yifan Wu, Yiyu Chen, Yenchi Tseng, Sicheng Liu, and Yuyu Luo. They really put forward a benchmark that shifts the focus from just generating pretty UIs to actually understanding the functional logic behind those UIs when you only have a picture of them.
Jane: It seems like the real implication here is that we need better ways to test if an AI has grasped the underlying state and flow of an application, not just its surface appearance.
Lu: I see this opening for really creative applications where we can build systems that require complex, multi-step interactions that are hard to prompt purely through language.
Meng: From a practical standpoint, this suggests that if we want AI agents to reliably navigate and operate software interfaces autonomously, they have to master this kind of inference.
Lalam: For me, the impact is about improving how AI interacts with our digital culture; if models can infer complex state transitions, those AI systems become much more capable of handling nuanced user flows in real-world applications.
Tom: Exactly. So, we're seeing that visual fidelity alone isn't a guarantee of functional understanding when it comes to interaction inference.
Jane: That’s the core message: we need metrics that specifically target how well an AI can reconstruct the application's behavior from just pixels, and UI2App is the first real tool for that job.
Lu: The way they structured those seven interaction categories provides a really useful roadmap for future work in building these specialized evaluation protocols.
Meng: I think we should be watching how this taxonomy helps guide how we design our own testing suites for complex software agents.
Lalam: I'm excited because it shows that the complexity of state management, specifically that cross-route persistence they highlighted, is a real barrier that models need to overcome.
Grace Man Chen, Litao Guo, Yifan Wu, Yiyu Chen, Yenchi Tseng, Sicheng Liu
The Hong Kong University of Science and Technology
cs.SE, cs.AI, cs.CV
Submitted: 2026-07-07
Updated: 2026-10-04
Project page: https://openai.com/index/gpt-5-4-thinking-system-card
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 90/100
The gist: Large language models have demonstrated growing competence in web page generation, but existing text-driven approaches are limited by complex prompts, and image-driven paradigms lack systematic
Key concepts
- Interaction Inference Scoring (IIS)
- This core metric quantifies a model's ability to predict and implement actions implied by a screenshot. It assesses coverage, result, and scope across three interaction types: UI state, data state, and cross-route persistence. A higher IIS means the model successfully infers the intended application behavior.
- Cross-Route State Persistence (S3)
- This refers to the most difficult type of interaction to infer—understanding how data or settings persist across different parts or routes of a web application. The paper identifies this as a major bottleneck, with half of tested models failing to correctly handle this complexity.
- Latent Affordance Inference
- This is the underlying capability being tested: the model's ability to understand hidden, implied behaviors (affordances) within a visual interface. The research suggests that this inference capability does not scale predictably with model size or image quality alone.
Terminology
Summary
Large language models have demonstrated growing competence in web page generation, but existing text-driven approaches are limited by complex prompts, and image-driven paradigms lack systematic evaluation of interaction capabilities. UI2App introduces a benchmark targeting interaction inference—the ability to recover application behavior from screenshots alone without textual guidance—to address this gap.
The gist
Inferring complete interaction behavior from static screenshots remains a key challenge for models, as visual fidelity does not imply interaction-inference capability, and cross-route state persistence is a frontier-wide bottleneck.
Benchmark Design and Evaluation Protocol
UI2App comprises 327 screenshots grouped into 45 state-coherent screenshot sets representing runnable multi-route web applications. The evaluation protocol assesses each artifact along four dimensions: executability (EXEC@1 / EXEC@3), navigation reachability (NRS), visual fidelity (VFS), and interaction inference (IIS). The core metric, IIS, is assessed by functional correctness and state-management complexity, utilizing an interaction taxonomy
of seven categories. This taxonomy organizes interactions into S1 UI-state, S2 data-state, and S3 cross-route persistence.
Interaction Inference Scoring (IIS)
The IIS quantifies how well a model infers and realizes screenshot-implied interactions by evaluating coverage, implementation result, and state-logic complexity across the seven categories. Each interaction is assessed along three axes: coverage (is the interaction implied?), result (is it fully realized, partially realized, or failed?), and scope (S1 UI-state, S2 data-state, S3 cross-route persistence). The model-level IIS is calculated using a formula that weights each scope level linearly with weights wS1=1, wS2=2, wS3=3. This taxonomy allows for structured assessment without requiring a unique ground-truth execution trace.
Evaluation Metrics and Findings
The four metrics are EXEC@k (Executability), NRS (Navigation Reachability Score), VFS (Visual Fidelity Score), and IIS. Experiments on six frontier vision-language models reveal a marked capability mismatch: the visual-fidelity leader scores only 7.5 on IIS, ranking fourth and trailing the IIS leader by 5.2×. High-complexity interactions, such as cross-page state (S3), remain a pervasive bottleneck, with half of the evaluated models scoring exactly zero on this dimension.
Failure Analysis and Scaling
Failure analysis across six frontier models reveals distinct mechanisms: EXEC@3 losses are dominated by scaffold-respect violations (e.g., C2: scaffold package.json overwritten) for GLM-4.6V, while GPT-5.4's losses are dominated by factuality errors (C1: hallucinated icon export). Furthermore, the within-family scaling analysis on the Qwen2.5-VL ladder shows a non-gradual phase transition between 32B and 72B for executability, indicating that parameter scaling alone does not saturate UI2App’s requirements. The paper concludes that latent-affordance inference is neither monotonic in model scale nor in visual fidelity, with IIS being the metric designed to expose this divergence.
Key Contributions
The contributions include: 1) Releasing UI2App, the first benchmark targeting interaction inference rather than specification-following from image-only input. 2) Introducing an end-to-end evaluation protocol spanning build executability, navigation reachability, visual fidelity, and interaction inference (IIS). 3) Establishing baselines for image-only interaction inference and demonstrating that visual fidelity does not imply interaction-inference capability. 4) Identifying cross-route state persistence as a frontier bottleneck.
Failure Taxonomy Details
The unrecovered EXEC@3 failures decompose into nine fine-grained modes, collapsing into five high-level causes: hallucinated imports (C1, C6), scaffold-file overwrites (C2, C8), undeclared / wrong-name source dependencies (C3), broken import paths (C4, C5), and parse-level or runtime errors (C7, C9). The Instruction-Compliance Score (ICS) shows that self-debug retries a failing app up to three rounds with the same model and a fixed feedback window. Models like GLM-4.6V show a 53% ICS, dominated by scaffold-respect violations. The analysis of case studies illustrates that IIS targets latent affordance inference, which is precisely the regime where pixel-fidelity scores cannot expose behavioral differences.
Case Studies on Latent Affordances
Case studies illustrate four claims: bimodal interaction emergence, domain-level behavioral inference, build-time failure modes, and VFS sub-metric divergence. For instance, in the 31 genshin-music task, only Kimi K2.5 instantiates a complete Web-Audio signal chain (audio runtime), while other models emit visually faithful but inert components.
Improvements for AI systems
Here are specific improvements to AI systems based on the UI2App benchmark research, detailing what these enhanced systems can achieve:
)1. Develop Interaction-Aware Generation
Models (IIS Focus):
The primary improvement is shifting model objective from mere visual fidelity (VFS) to functional correctness and state management.
-
Specific Improvement: Fine-tune or architect models using the UI2App taxonomy (7 interaction categories across 3 complexity scopes: S1, S2, S3). Implement reward functions that explicitly penalize
frozen façade
errors (e.g., failing to implement cross-route state persistence, C07) more heavily than simple visual mismatches. -
What the improved system can do: It will generate web applications where user actions (like adding an item to a cart and navigating to checkout) result in a logically consistent, stateful transition across multiple routes, rather than just rendering static snapshots of each route individually.
)2. Implement Robust Self-Debug and Iterative Repair Pipelines (EXEC@3 Focus):
The system must be designed not just for one-shot generation, but for resilient self-correction during the build process.
-
Specific Improvement: Integrate the three-stage self-debug protocol (File Localization Fallback, Build Error Repair, Runtime Error Repair). This requires the model to ingest build logs and file contents and perform targeted edits rather than generating a complete rewrite from scratch upon failure. The system must be trained to diagnose specific error modes (e.g., C1 hallucinated imports vs. C2 scaffold overwrites).
-
What the improved system can do: It will achieve significantly higher executability (EXEC@3) by reliably fixing the root cause of build failures—whether it's a missing dependency declaration, an incorrect import path, or an instruction-compliance violation—allowing it to move from failed builds to working artifacts much faster.
)3. Enhance Domain-Specific Behavioral Inference (Latent Affordance Detection):
The system needs to learn what the UI implies
beyond what is explicitly coded or visually rendered.
-
Specific Improvement: Train models on the case studies (e.g., 31 genshin-music, 67 shadboard) where crucial behaviors (audio synthesis, answer validation logic) are only implied by sparse visual cues (layout, toolbars). The training should focus on mapping low-fidelity visual features to high-level functional affordances.
-
What the improved system can do: It will move beyond surface mimicry to infer deep, latent affordances—such as recognizing a visually plausible
drop zone
that implies an HTML5 drag-and-drop handler or inferring acheck button
that triggers complex validation logic—even when those handlers are entirely absent from the screenshot.
)4. Decouple Visual Fidelity from Functional Realization:
The architecture must prevent models from conflating high visual scores with high functional scores.
-
Specific Improvement: Use the conditional metrics (VFS⋆ and IIS⋆) as primary evaluation criteria during training, rather than relying solely on zero-imputation headline scores. This forces the model to optimize for correctness even when it achieves a visually
good
but functionally inert output. -
What the improved system can do: It will produce artifacts that are visually indistinguishable from high-quality mockups (high VFS) but are functionally unusable because they lack necessary state management or interaction logic (low IIS). This provides a crucial diagnostic signal for developers to understand where visual training needs to be supplemented with behavioral supervision.
)5. Implement State-Space Reasoning for Complex Applications:
For multi-route applications, the model must maintain a mental model of global state across navigation boundaries.
-
Specific Improvement: Introduce structured planning steps that explicitly track
state variables
(e.g., Cart contents, User Auth status) as they transition between routes. The planning prompt should require the model to define and manage these state variables across multiple files/routes in its JSON plan before generation begins. -
What the improved system can do: It will reliably handle complex, multi-page workflows (like e-commerce checkouts or dashboard navigation) where data synchronization must survive route changes, ensuring that a cart added on one page is correctly reflected and persistent on another.
Abstract
Large language models (LLMs) have demonstrated growing competence in generating web pages from UI screenshots, which convey both visual structure and cues to application behavior. Yet most screenshot-to-code benchmarks emphasize visual fidelity, while interactive generation benchmarks often supply behavioral specifications or demonstrated transitions. Whether models can infer and realize interactions from static screenshots alone remains insufficiently evaluated. We introduce UI2App to evaluate interaction inference: inferring and realizing application behavior from static visual cues without added behavioral guidance. UI2App comprises 600 screenshots organized into 95 state-coherent sets for runnable multi-route web applications. Our end-to-end pipeline evaluates each artifact along three dimensions: executability, visual fidelity, and interaction inference. The interaction metric (IIS) assesses functional correctness and state-management complexity, crediting valid implementations rather than requiring a match to a single reference. Experiments on six frontier vision-language models reveal a marked mismatch between visual fidelity and interaction realization: the visual-fidelity leader scores only 8.1 on IIS, ranking fourth, while the IIS leader achieves 4.5 times that score. High-complexity interactions such as cross-route state persistence remain a major bottleneck, with five of the six models scoring at most 3.5 on this dimension. Overall, these results highlight interaction inference as a key challenge in generating functional web applications from static screenshots.
Sources
- IWR-Bench: Can LVLMs reconstruct interactive webpage from a user interaction video?
- PSD2Code: Automated Front-End Code Generation from Design Files via Multimodal Large Language Models
- DesignCoder: Hierarchy-Aware and Self-Correcting UI Code Generation with Large Language Models
- WebVR: Benchmarking Multimodal LLMs for WebPage Recreation from Videos via Human-Aligned Visual Rubrics
- Advancing vision-language models in front-end development via data synthesis
- DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Vision2Web: A Hierarchical Benchmark for Visual Website Development with Agent Verification
- Kimi K2.5: Visual Agentic Intelligence
- Unlocking the conversion of Web Screenshots into HTML Code with the WebSight Dataset
- Code Llama: Open Foundation Models for Code
- MRWeb: An Exploration of Generating Multi-Page Resource-Aware Web Code from UI Designs
- VSA:Visual-Structural Alignment for UI-to-Code
- DesignBench: A Comprehensive Benchmark for MLLM-based Front-end Code Generation
- ArtifactsBench: Bridging the Visual-Interactive Gap in LLM Code Generation Evaluation
- Widget2Code: From Visual Widgets to UI Code via Multimodal LLMs
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties