Looks Right, Works Right: A Project-Level Benchmark for Multi-Screen Mobile App Generation
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Looks Right, Works Right: A Project-Level Benchmark for Multi-Screen Mobile App Generation".
Jane: The paper was written by N/A - Authors not found in provided context. from.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Paper discussion segment 1: Tom: Continuing our discussion on "Looks Right, Works Right: A Project-Level Benchmark for Multi-Screen Mobile App Generation," we previously established that the benchmark is about moving beyond isolated testing units. Now, Jane, can you summarize what the paper suggests regarding the actual scope of testing that needs to be covered?
Jane: The authors propose a significant conceptual shift by arguing that AI must be tested using real-world data—data sets that are deliberately messy or imperfect. They aren't looking for success; they're asking the AI to model failure gracefully, which is a huge departure from traditional testing models.
Lu: That’s a crucial detail, Tom. It means the AI isn't just proving it *can* build a beautiful interface based on perfect mockups; it’s actually proving it can build an interface that *survives* the chaos inherent in human input and legacy systems, which is far more realistic.
Meng: From an engineering standpoint, this goes deeper than just data messiness. They are also suggesting that the benchmark must test standardized component registries linked to real-time APIs. It needs to validate the entire integration layer, not just the visual layout of widgets on a screen.
Lalam: And on top of that, we have to consider global usability standards for this system. The authors stress that a functional benchmark must include tests for diverse user needs—things like different screen sizes, contrasting color schemes, or supporting non-Latin writing systems globally.
Tom: So it’s not just about the technical robustness of the backend; it's about the human robustness across diverse contexts and cultures. Jane, how do these three elements—messy data, APIs, and global usability—all connect to redefine testing?
Jane: Precisely. The improvement isn't merely adding more tests; it’s redefining *what* constitutes a valid test—it must encompass data integrity, system interoperability, and diverse user needs all at once. Next, we’ll dive into the specific actionable improvements they suggest for making this benchmark even stronger in practice.
Paper discussion segment 2: Jane: Picking up our conversation on "Looks Right, Works Right: A Project-Level Benchmark for Multi-Screen Mobile App Generation," we’ve established that the benchmark needs to test data messiness and global usability. Now, let's focus on the specific improvements the authors suggest—the actionable steps to move this from a theoretical concept into an industry standard.
Tom: One of the most significant suggestions is moving beyond clean, curated data sets. They argue that for AI to be truly useful in production environments, it must be tested using messy, real-world data—data that has inconsistencies or missing fields. This forces the generative model to prove its resilience under imperfect conditions.
Lu: That moves the AI from being a mere designer into being a constrained system engineer. It has to balance optimal functionality with practical deployment realities, which is significantly harder than just designing for the best-case scenario we usually build for in academia environments.
Meng: And tying into that idea of constraint, they also suggest creating standardized component registries linked to real-time APIs. This means the benchmark needs to simulate complex data fetching—like pulling user data from three different backend services simultaneously—and test the *integration* layer itself under pressure.
Lalam: I think we need to pay close attention here because this touches on global accessibility and internationalization standards that are critical for modern apps. The authors stress that a functional benchmark must include tests for diverse user needs, such as testing the app's behavior with contrasting color schemes or in non-Latin writing systems for global usability.
Jane: Exactly. So, the improvement isn't just adding more tests; it’s changing *what* we consider a valid test—it has to encompass data integrity, system interoperability, and diverse human needs simultaneously for a product to pass.
Tom: The ultimate goal they paint is an AI that acts as a full-stack QA architect, anticipating every point of failure before a single developer has written the final line of code. This level of foresight is revolutionary. We'll discuss what this means for the future development lifecycle next.
Paper discussion segment 3: Jane: Picking up our conversation on "Looks Right, Works Right: A Project-Level Benchmark for Multi-Screen Mobile App Generation," we’ve established that the benchmark needs to test data messiness and global usability. Now, let's focus on the specific improvements the authors suggest—the actionable steps to move this from a theoretical concept into an industry standard.
Tom: One of the most significant suggestions is moving beyond clean, curated data sets. They argue that for AI to be truly useful in production environments, it must be tested using messy, real-world data—data that has inconsistencies or missing fields. This forces the generative model to prove its resilience under imperfect conditions rather than just perfect inputs.
Lu: That move fundamentally shifts the focus of evaluation; we are no longer testing for ideal output but for graceful failure handling, which is a much more useful skill in a commercial product lifecycle.
Meng: And tying into that idea of constraint, they also suggest creating standardized component registries linked to real-time APIs. This means the benchmark needs to simulate complex data fetching—like pulling user data from three different backend services simultaneously—and test the *integration* layer itself rigorously.
Lalam: I think we need to remember
Conclusion: Tom: So, if we pull together everything we discussed today about "Looks Right, Works Right: A Project-Level Benchmark for Multi-Screen Mobile App Generation," the core message is that assessing AI needs to move past checking isolated pieces of code and instead validate entire user journeys across a whole application.
Jane: Exactly. It represents a massive conceptual shift in how we view generative capability; we are now looking at underlying system architecture, not just the surface polish of an individual screen or component.
Lu: What this really implies is that AI must transition from being just a writing tool to acting as a fundamental systems architect, someone who anticipates failure points across the user's entire goal state.
Meng: I agree with Lu; it underscores that solving these project-level integration issues demands not only better generative models but also standardized tooling—like component registries—to manage the complexity reliably.
Lalam: And I keep coming back to how critical the human element is here; a truly functional benchmark has to prove usability for people using vastly different devices, languages, and accessibility needs globally.
Tom: So we are moving toward an expectation that AI can function as a full-stack QA architect right out of the gate. Jane?
Jane: It raises the bar significantly for every developer who relies on these tools; they now need to think about interoperability and failure modes from day one, not as an afterthought.
Tom: This entire discussion highlights how vital comprehensive testing is for building trustworthy digital products. We covered a lot of ground today regarding "Looks Right, Works Right: A Project-Level Benchmark for Multi-Screen Mobile App Generation."
Jane: It sets a much higher standard for what we expect from the next generation of these powerful tools.
Tom: And that leaves us ready to look at how these new standards will reshape the actual development lifecycle. Next up, we'll talk about resource allocation in modern software teams.
N/A - Authors not found in provided context.
cs.HC, cs.AI
Submitted: 2026-08-21
Updated: 2026-08-24
Code: https://github.com/anoa12159-hue/mobileforge_eval
Importance score: 78/100
The gist: The paper presents "Looks Right, Works Right: A Project-Level Benchmark for Multi-Screen Mobile App Generation," detailing a comprehensive evaluation framework designed to test the reliability and
Key concepts
- Project-Level Benchmark
- The paper proposes moving beyond testing individual components or screens. Instead, it requires validating the entire application's functionality across a whole user journey to ensure the AI can build a complete, robust product.
- Messy Real-World Data
- AI must be tested using data that is imperfect, inconsistent, or missing fields. This forces generative models to prove their resilience and handle failure gracefully, rather than just succeeding with perfect inputs.
- System Interoperability
- This concept requires the benchmark to test how different parts of an app work together. It involves validating the entire integration layer by simulating complex data fetching from multiple backend services or APIs.
- Global Usability
- A functional benchmark must account for diverse user needs worldwide. This includes testing for different screen sizes, contrasting color schemes, and supporting non-Latin writing systems to ensure accessibility across cultures.
Terminology
Summary
The paper presents Looks Right, Works Right: A Project-Level Benchmark for Multi-Screen Mobile App Generation,
detailing a comprehensive evaluation framework designed to test the reliability and structural integrity of large language models (LLMs) in generating complex, multi-screen mobile applications.
Methodologically, the research employs an annotation pipeline that involves human review across various components. This process allows reviewers to inspect each page’s type, description, key elements, layout, and primary action,
with the ability to pass, edit, or reject auto-drafted entries (Figure 4). Furthermore, relationship testing is critical; the system surfaces Parent–child and tab pairs with source/target screenshots and the trigger description
for review (Figure 5).
The benchmark assesses performance across several key structural failure modes. The annotation pipeline quality is measured using metrics such as Recall: fraction of final reviewed items already produced by the autodrafting model,
Precision: fraction of auto-drafted items retained in the reviewed set,
and Unchanged rate: fraction of auto items retained with no field-level edits
(Table 9).
The study categorizes and analyzes four distinct types of failure, which are critical for evaluating functional correctness:
-
C1: Blank starting page: Occurs when
The source page renders empty or as a stub placeholder with no interaction target, so no navigation action can be dispatched
(Figure 7). -
C2: Route mapping error: Occurs when the agent declares a route, but
the URL resolves to a different page than the one referenced by the test; the model has correctly produced the page, but the global routing table maps the navigation action to the wrong page
(Figure 9). -
C3: Target unreachable or occluded: This failure is characterized by a malformed page where
The action target is not present in the rendered viewport, either because the layout omits it entirely or because another element covers it.
The paper notes this isstructurally distinct from C4
(Figure 10). -
C4: Target clickable but unresponsive: This failure occurs when
The page renders, the affordance is in the right place, but the click registers no effect, consistent with a missing or mis-wired onClick handler
(Figure 13).
Model performance is evaluated across these failure types using dedicated hand-labelled interactive-navigation samples (Table 10). The analysis of model capabilities reveals that color consistency has saturated across all frontier models (0.74–0.85, 15% spread), while structural indicators (LoC, LoC/file, reuse, dead-component rate) differentiate sharply
(Table 12).
In addition to functional failures, the paper also tracks the composition of context tokens per model across different inputs—including system prompt plus initial user text,
write file echo: write file tool results returned to the agent as context, and various other tool results—to understand how models account for image content and structured data (Table 8).
Overall, the research provides detailed metrics on code maintainability indicators, such as Files: number of source files,
Shared comp.: count of components under src/components/,
and LoC s.d.: standard deviation of LoC across apps for this model
(Table 12), allowing for a granular comparison among models like Claude Opus 4.6, GPT-5, Gemini 2.5 Pro, and others in the context of generating robust mobile applications.
Improvements for AI systems
Based on the detailed analysis of failure modes, context composition, and structural indicators presented in these tables, the current generation of LLM-based agents requires fundamental architectural upgrades. The core deficiency is not semantic understanding but structural, functional reliability within a constrained execution environment (the application UI).
Here are three major areas for improvement, detailing the mechanism and the resulting capabilities of the enhanced AI system.
The current models treat navigation as a sequence of text predictions, which is why they fail at route mapping (C2), identifying occluded targets (C3), or recognizing inert elements (C4). We must move beyond simple visual or textual grounding.
Mechanism:
-
DOM/API Schema Integration: The model must be provided with a real-time, structured representation of the target application's Document Object Model (DOM) and its API routing table before generating an action.
-
Constraint Layering: Any proposed action (e.g.,
click,navigate) must pass through a validation layer that checks:
-
Existence: Does the element ID/XPath exist in the current DOM state? (Mitigates C3).
-
Affordance: Is the element marked as interactive (
role="button",onClickhandler present) and not visually obscured by a high-Z-index overlay? (Mitigates C4). -
Route Mapping: Does the proposed action ID map to a valid, canonical route in the global routing table? (Mitigates C2).
What the Improved AI System Can Do:
-
Guaranteed Functional Navigation: The system will never propose a navigation action that is structurally impossible or functionally dead. If an element is visually present but has no associated
onClickhandler (C4), the SGAE flags it as an error before execution, reportingAffordance Failure: Element exists but lacks functional binding.
-
Precision in State Transitions: It can reliably predict the next valid state of the application, drastically improving Recall and Precision by ensuring that every generated step is guaranteed to transition to a stable, reachable page state.
The current context composition metrics show that tool results (read file, run command) are appended as flat text chunks. This prevents the model from understanding the relationship between the failure source, the attempted action, and the resulting context change.
The analysis shows that structural indicators (LoC, shared components) differentiate models. The current agents are treating the UI as a monolithic canvas. We must train them to reason at the level of reusable code components.
Sources
- SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
- PairBench: Are Vision-Language Models Reliable at Comparing What They See?
- Program Synthesis with Large Language Models
- Advancing vision-language models in front-end development via data synthesis
- Evaluating Large Language Models Trained on Code
- ScreenCoder: Advancing Visual-to-Code Generation for Front-End Automation via Modular Multimodal Agents
- Unlocking the conversion of Web Screenshots into HTML Code with the WebSight Dataset
- SWE-Bench Mobile: Can Large Language Model Agents Develop Industry-Level Mobile Applications?
- Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception
- OmniParser for Pure Vision Based GUI Agent
- UI2Code^N: UI-to-Code Generation as Interactive Visual Optimization
Related papers
- EduGage: A Multimodal Dataset and Benchmark for Sensor-Based Momentary Assessment of Engagement in Self-Guided Video Learning
- EvoDesign: Agentic Editable Diagram Creation via Design Expertise Evolution
- HAGI++: Head-Assisted Gaze Imputation and Generation
- Linking Behaviour and Perception to Evaluate Meaningful Human Control over Partially Automated Driving
- Review of Explainable Decision Support and Adaptive Human-Machine Interfaces for Automation Transparency in Maritime Autonomous Surface Ships
- Towards Cognitive Process-Aware Proactive Writing Support