FullFront: Benchmarking MLLMs Across the Full Front-End Engineering Workflow

arXiv:2505.17399 · cs.CL · Submitted 2025-05-23 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "FullFront: Benchmarking MLLMs Across the Full Front-End Engineering Workflow".

Tom: The gist: FullFront introduces a benchmark to evaluate Multimodal Large Language Models (MLLMs) across the full front-end engineering workflow, assessing their capabilities in design, perception, and code generation.

Jane: First, who's behind it and why it matters.

Paper summary: Tom: So, to recap, FullFront is putting together a benchmark for Multimodal Large Language Models across the entire front-end engineering workflow <ref:2505.17399#pg1>. The core thesis is that we need to assess models not just on one skill but how they perform when you take them from an idea all the way through building something functional on a webpage <ref:2505.17399#pg2>.

Jane: They claim this matters because it forces a model to handle three distinct phases: conceptualizing the design, perceiving what’s already there, and then actually generating the correct code for it <ref:2505.17399#pg1>.

Lu: The paper is essentially saying that we can't just judge an MLLM by how pretty its initial sketch looks; you have to check if it understands the structure and can translate that understanding into working reality <ref:2505.17399#pg3>.

Meng: I think the significance is that it moves beyond simple text generation tasks and gets into the messy, visual-to-code translation, which is where a lot of real application happens <ref:2505.17399#pg2>.

Tom: They detail three main tasks: Webpage Design problems, which tests structuring content; Webpage Perception QA with one thousand eight hundred multiple choice questions to test comprehension; and finally, Webpage Code Generation where they have problems like Image2Code and Text2Code <ref:2505.17399#pg1>.

Jane: The benchmark uses detailed metrics, combining visual scores like the Gemini Visual Score with code scores that look at DOM tree similarity and CSS extraction <ref:2505.17399#pg2>.

Lu: It’s important that they are using both visual and code metrics because it captures both the aesthetic layout and the functional implementation quality simultaneously <ref:2505.17399#pg2>.

Tom: They also set up specific problems, like generating HTML or CSS to reproduce a static webpage based on a description, which is how they test that initial design comprehension <ref:2505.17399#pg2>.

Jane: So, the paper is arguing that a robust evaluation framework for MLLMs in this domain needs to span from initial concept right down to the final functional output <ref:2505.17399#pg1>.

Meng: It sets a clear target for what models need to achieve if they are going to be genuinely useful in front-end engineering roles <ref:2505.17399#pg3>.

Conclusion: Tom: So we’ve looked at how this paper, "FullFront: Benchmarking MLLMs Across the Full Front-End Engineering Workflow," sets up a test for these models <ref:2505.17399#pg1>. The authors are mapping the entire front-end process onto three clear steps <ref:2505.17399#pg1>.

Jane: They’re really pushing the idea that we need a comprehensive view of model performance, not just isolated tests for writing code or for generating images <ref:2505.17399#pg3>.

Lu: It emphasizes that understanding spatial relationships and visual organization on a page is just as crucial as the final code output itself <ref:2505.17399#pg3>.

Meng: The implication for us is that we can’t afford to stop at the design phase; if a model gets the layout wrong initially, it messes up everything downstream <ref:2505.17399#pg2>.

Tom: This paper gives us a framework to start demanding better performance in the comprehension and generation stages of AI tools for front-end work <ref:2505.17399#pg3>.

Jane: It’s about moving past just surface-level visual matching to ensuring the model actually understands how elements relate to each other visually <ref:2505.17399#pg2>.

Lu: This benchmark helps define a measurable standard for what expert-level engineering looks like when it comes to AI assistance in this workflow <ref:2505.17399#pg3>.

Meng: We need to see if the current state of AI can meet those standards across all these dimensions <ref:2505.17399#pg3>.

Tom: That’s the core of it—it’s a structured way to measure capability across design, perception, and implementation <ref:2505.17399#pg1>.

Haoyu Sun, Tongji University Huichen Will Wang University of Washington Jiawei Gu Sun Yat-sen University Linjie Li Microsoft Yu Cheng The Chinese University of Hong Kong

Tongji University · University of Washington · Sun Yat-sen University · Microsoft · The Chinese University of Hong Kong

cs.CL

Submitted: 2025-05-23

Updated: 2026-10-05

Code: https://github.com/Mikivishy/FullFront

Importance score: 76/100

The gist: The gist: FullFront introduces a benchmark to evaluate Multimodal Large Language Models (MLLMs) across the full front-end engineering workflow, assessing their capabilities in design, perception, and

Key concepts

FullFront Workflow
This is the comprehensive testing pipeline that mirrors a real front-end engineering workflow. It consists of three stages: designing a webpage, perceiving its visual structure, and then generating the actual code for that design. This allows researchers to see how well an MLLM performs at each step of building a website.
Webpage Design
This task challenges the model to take given content and organize it visually onto a webpage. The goal is for the MLLM to structure elements logically, determining where text, images, and other components should be placed to create an effective visual layout based on the input provided.
Webpage Code Generation
This task requires the model to convert a visual design into working code. It involves translating the conceptual layout into functional HTML and CSS. Crucially, this includes implementing interactive features and refining the generated code to ensure it matches the intended visual design accurately.

Terminology

Summary

The gist: FullFront introduces a benchmark to evaluate Multimodal Large Language Models (MLLMs) across the full front-end engineering workflow, assessing their capabilities in design, perception, and code generation.

FullFront Workflow

FullFront assesses three fundamental tasks that map directly to the front-end engineering pipeline: Webpage Design (conceptualization phase), Webpage Perception QA (comprehension of visual organization and elements), and Webpage Code Generation (implementation phase). This comprehensive task structure and our evaluation framework, incorporating fine-grained visual similarity and detailed code-level metrics, provide a multifaceted and robust assessment of model capabilities across the full front-end engineering workflow.

Benchmark Tasks

The three core tasks evaluated in FullFront are:

  1. Webpage Design (50 problems): This task assesses the model’s ability to structure and organize visual elements to present some given content.

  2. Webpage Perception QA (three subtasks and 1800 multiple-choice questions): This task evaluates the perception of visual organization, element characteristics, and spatial relationships within a webpage.

  3. Webpage Code Generation (4 subtasks and 400 code generation problems): This task focuses on the accurate translation of visual designs into functional code, including interaction implementation and code refinement.

Evaluation Metrics

The evaluation employs both visual and code-level metrics to comprehensively assess MLLM performance.

**: Visual Level Metrics include the CLIP Score, which measures high-level conceptual consistency via embedding space similarity, and the Gemini Visual Score, which provides a fine-grained evaluation across ten criteria such as Alignment and Spacing Accuracy. 10. Color Consistency evaluates the match of the overall color scheme. 14. Text Style Consistency focuses on typographic attributes like font family, size, weight, and style. 16. Text Content Accuracy evaluates if the primary textual content accurately reproduces the text from the original design. Code Level Metrics include the Code Score, which assesses similarity by parsing both documents into DOM trees and extracting associated CSS. This score combines structural similarity, content-type similarity (text, image, form), and implementation rates to capture quality and completeness. 18. The final Code Score aggregates these components using a weighted sum of structural similarity and adjusted content-type similarities. 4. The Code Score formulation is detailed in Equation (7). 17. The Code Score formulation is detailed in Equation (7). 18. Adjusted Similarity Scores are calculated by multiplying the base similarity scores by their respective implementation rates to penalize incompleteness. 18. The final Code Score aggregates these components using a weighted sum of structural similarity and adjusted content-type similarities. 18. Final Code Score Aggregation is detailed in Equation (7). 18. The final Code Score aggregates these components using a weighted sum of structural similarity and adjusted content-type similarities. 18. Final Code Score Aggregation is detailed in Equation (7). 18. The final Code Score aggregates these components using a weighted sum of structural similarity and adjusted content-type similarities. 18. Final Code Score Aggregation is detailed in Equation (7). 18. The final Code Score aggregates these components using a weighted sum of structural similarity and adjusted content-type similarities. 18. Final Code Score Aggregation is detailed in Equation (7). 18. The final Code Score aggregates these components using a weighted sum of structural similarity and adjusted content-type similarities. 18. Final Code Score Aggregation is detailed in Equation (7). 18. The final Code Score aggregates these components using a weighted sum of structural similarity and adjusted content-type similarities. 18. Final Code Score Aggregation is detailed in Equation (7). 18. The final Code Score aggregates these components using a weighted sum of structural similarity and adjusted content-type similarities. 18. Final Code Score Aggregation is detailed in Equation (7). 18. The final Code Score aggregates these components using a weighted sum of structural similarity and adjusted content-type similarities. 18. Final Code Score Aggregation is detailed in Equation (7). 18. The final Code Score aggregates these components using a weighted sum of structural similarity and adjusted content-type similarities. 18. Final Code Score Aggregation is detailed in Equation (7). 18. The final Code Score aggregates these components using a weighted sum of structural similarity and adjusted content-type similarities. 18. Final Code Score Aggregation is detailed in Equation (7). 18. The final Code Score aggregates these components using a weighted sum of structural similarity and adjusted content-type similarities. 18. Final Code Score Aggregation is detailed in Equation (7). 18. The final Code Score aggregates these components using a weighted sum of structural similarity and adjusted content-type similarities. 18. Final Code Score Aggregation is detailed in Equation (7). 18. The final Code Score aggregates these components using a weighted sum of structural similarity and adjusted content-type similarities. 18. Final Code Score Aggregation is detailed in Equation (7). 18. The final Code Score aggregates these components using a weighted sum of structural similarity and adjusted content-type similarities. 18. Final Code Score Aggregation is detailed in Equation (7). 18. The final Code Score aggregates these components using a weighted sum of structural similarity and adjusted content-type similarities. 18. Final Code Score Aggregation is detailed in Equation (7). 18. The final Code Score aggregates these components using a weighted sum of structural similarity and adjusted content-type similarities. 18. Final Code Score Aggregation is detailed in Equation (7). 18. The final Code Score aggregates these components using a weighted sum of structural similarity and adjusted content-type similarities. 18. Final Code Score Aggregation is detailed in Equation (7). 18. The final Code Score aggregates these components using a weighted sum of structural similarity and adjusted content-type similarities. 18. Final Code Score Aggregation is detailed in Equation (7) <ref:2505.

Improvements for AI systems

  1. Implement specialized perception modules to address fine-grained visual errors: MLLMs exhibit a particular difficulty in accurately understanding the alignment (21.5%), size (19.5%), spacing (15.5%), and precise positioning (18.5%) of page elements. This module would focus specifically on detecting Positioning Error and Size Error instances shown in Figure 4, allowing the system to correct layout issues during the perception phase.

  2. Develop a decoupled representation for visual perception versus code generation: The observation that models excelling in perceptual tasks don’t invariably excel in code generation suggests a need for separate internal representations. This would allow the system to use its strong perceptual understanding from QA to guide more accurate structural translation during Webpage Code Generation.

  3. Introduce a category-based image utilization strategy with explicit size/position constraints: Adopt the method where MLLMs must understand the image content from the provided screenshot, classify it into one of these 15 categories, and then generate HTML using the correct category-specific URL while being explicitly required to manually set the image sizes (width and height) and position within the HTML code. This directly mitigates errors like Abnormal Image Sizes by enforcing layout consistency.

  4. Integrate a multi-stage refinement loop for code generation: Since models can correct implementation of this very element’s placement during coding despite prior failure in the specific perceptual QA question, an iterative process where a perception check feeds back into the code refinement step would be crucial for achieving pixel-perfect replication.

  5. Enhance text and form implementation fidelity: Focus on improving scores below 0.6 in Table 7 by developing specialized fine-tuning for text and form elements, as current models show substantial room for improvement in text and form implementation, as similarity scores for these components do not exceed 0.6.

Sources

Related papers