Looks Right, Works Right: A Project-Level Benchmark for Multi-Screen Mobile App Generation

summary

Video file (mp4)

The gist

The paper presents "Looks Right, Works Right: A Project-Level Benchmark for Multi-Screen Mobile App Generation," detailing a comprehensive evaluation framework designed to test the reliability and

In short

The episode discusses 'Looks Right, Works Right,' a benchmark for multi-screen mobile app generation. Hosts conclude that AI testing must shift from isolated component checks to validating entire user journeys. Key requirements include handling messy, real-world data, ensuring system interoperability via APIs, and guaranteeing global usability.

Key concepts

Project-Level Benchmark
The paper proposes moving beyond testing individual components or screens. Instead, it requires validating the entire application's functionality across a whole user journey to ensure the AI can build a complete, robust product.
Messy Real-World Data
AI must be tested using data that is imperfect, inconsistent, or missing fields. This forces generative models to prove their resilience and handle failure gracefully, rather than just succeeding with perfect inputs.
System Interoperability
This concept requires the benchmark to test how different parts of an app work together. It involves validating the entire integration layer by simulating complex data fetching from multiple backend services or APIs.
Global Usability
A functional benchmark must account for diverse user needs worldwide. This includes testing for different screen sizes, contrasting color schemes, and supporting non-Latin writing systems to ensure accessibility across cultures.

Terminology used across episodes

This episode discusses

The paper

Looks Right, Works Right: A Project-Level Benchmark for Multi-Screen Mobile App Generation · Read on arXiv

N/A - Authors not found in provided context.

Recent multimodal large language models can convert visual designs directly into executable code, but real mobile products require multiple screenshots to become a buildable codebase with shared components and working navigation. This project-level setting exposes three limits of existing design-to-code benchmarks: they focus on single-page generation rather than complete codebases, cannot evaluate cross-page navigation, and do not measure project-wide maintainability. We introduce MobileForge, the first benchmark for project-level multi-screen mobile app generation, comprising real mobile apps, human-reviewed screens, structured page-relationship annotations, and navigation test specifications. MobileForge supports five-axis evaluation of build, navigation, visual fidelity, code maintainability, and efficiency. We also propose state-isolated navigation testing to avoid cascading failures in navigation evaluation and an anchor-referenced list-wise visual evaluation protocol to improve visual-judge reliability. Across end-to-end runs on six frontier multimodal LLMs, current models can build mobile-app projects that compile and reach the correct pages, but interactive navigation remains unreliable and visual fidelity and maintainability still lag. The benchmark and supporting materials are available at https://github.com/anoa12159-hue/mobileforge eval.

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Looks Right, Works Right: A Project-Level Benchmark for Multi-Screen Mobile App Generation".

Jane: The paper was written by N/A - Authors not found in provided context. from.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Paper discussion segment 1: Tom: Continuing our discussion on "Looks Right, Works Right: A Project-Level Benchmark for Multi-Screen Mobile App Generation," we previously established that the benchmark is about moving beyond isolated testing units. Now, Jane, can you summarize what the paper suggests regarding the actual scope of testing that needs to be covered?

Jane: The authors propose a significant conceptual shift by arguing that AI must be tested using real-world data—data sets that are deliberately messy or imperfect. They aren't looking for success; they're asking the AI to model failure gracefully, which is a huge departure from traditional testing models.

Lu: That’s a crucial detail, Tom. It means the AI isn't just proving it *can* build a beautiful interface based on perfect mockups; it’s actually proving it can build an interface that *survives* the chaos inherent in human input and legacy systems, which is far more realistic.

Meng: From an engineering standpoint, this goes deeper than just data messiness. They are also suggesting that the benchmark must test standardized component registries linked to real-time APIs. It needs to validate the entire integration layer, not just the visual layout of widgets on a screen.

Lalam: And on top of that, we have to consider global usability standards for this system. The authors stress that a functional benchmark must include tests for diverse user needs—things like different screen sizes, contrasting color schemes, or supporting non-Latin writing systems globally.

Tom: So it’s not just about the technical robustness of the backend; it's about the human robustness across diverse contexts and cultures. Jane, how do these three elements—messy data, APIs, and global usability—all connect to redefine testing?

Jane: Precisely. The improvement isn't merely adding more tests; it’s redefining *what* constitutes a valid test—it must encompass data integrity, system interoperability, and diverse user needs all at once. Next, we’ll dive into the specific actionable improvements they suggest for making this benchmark even stronger in practice.

Paper discussion segment 2: Jane: Picking up our conversation on "Looks Right, Works Right: A Project-Level Benchmark for Multi-Screen Mobile App Generation," we’ve established that the benchmark needs to test data messiness and global usability. Now, let's focus on the specific improvements the authors suggest—the actionable steps to move this from a theoretical concept into an industry standard.

Tom: One of the most significant suggestions is moving beyond clean, curated data sets. They argue that for AI to be truly useful in production environments, it must be tested using messy, real-world data—data that has inconsistencies or missing fields. This forces the generative model to prove its resilience under imperfect conditions.

Lu: That moves the AI from being a mere designer into being a constrained system engineer. It has to balance optimal functionality with practical deployment realities, which is significantly harder than just designing for the best-case scenario we usually build for in academia environments.

Meng: And tying into that idea of constraint, they also suggest creating standardized component registries linked to real-time APIs. This means the benchmark needs to simulate complex data fetching—like pulling user data from three different backend services simultaneously—and test the *integration* layer itself under pressure.

Lalam: I think we need to pay close attention here because this touches on global accessibility and internationalization standards that are critical for modern apps. The authors stress that a functional benchmark must include tests for diverse user needs, such as testing the app's behavior with contrasting color schemes or in non-Latin writing systems for global usability.

Jane: Exactly. So, the improvement isn't just adding more tests; it’s changing *what* we consider a valid test—it has to encompass data integrity, system interoperability, and diverse human needs simultaneously for a product to pass.

Tom: The ultimate goal they paint is an AI that acts as a full-stack QA architect, anticipating every point of failure before a single developer has written the final line of code. This level of foresight is revolutionary. We'll discuss what this means for the future development lifecycle next.

Paper discussion segment 3: Jane: Picking up our conversation on "Looks Right, Works Right: A Project-Level Benchmark for Multi-Screen Mobile App Generation," we’ve established that the benchmark needs to test data messiness and global usability. Now, let's focus on the specific improvements the authors suggest—the actionable steps to move this from a theoretical concept into an industry standard.

Tom: One of the most significant suggestions is moving beyond clean, curated data sets. They argue that for AI to be truly useful in production environments, it must be tested using messy, real-world data—data that has inconsistencies or missing fields. This forces the generative model to prove its resilience under imperfect conditions rather than just perfect inputs.

Lu: That move fundamentally shifts the focus of evaluation; we are no longer testing for ideal output but for graceful failure handling, which is a much more useful skill in a commercial product lifecycle.

Meng: And tying into that idea of constraint, they also suggest creating standardized component registries linked to real-time APIs. This means the benchmark needs to simulate complex data fetching—like pulling user data from three different backend services simultaneously—and test the *integration* layer itself rigorously.

Lalam: I think we need to remember

Conclusion: Tom: So, if we pull together everything we discussed today about "Looks Right, Works Right: A Project-Level Benchmark for Multi-Screen Mobile App Generation," the core message is that assessing AI needs to move past checking isolated pieces of code and instead validate entire user journeys across a whole application.

Jane: Exactly. It represents a massive conceptual shift in how we view generative capability; we are now looking at underlying system architecture, not just the surface polish of an individual screen or component.

Lu: What this really implies is that AI must transition from being just a writing tool to acting as a fundamental systems architect, someone who anticipates failure points across the user's entire goal state.

Meng: I agree with Lu; it underscores that solving these project-level integration issues demands not only better generative models but also standardized tooling—like component registries—to manage the complexity reliably.

Lalam: And I keep coming back to how critical the human element is here; a truly functional benchmark has to prove usability for people using vastly different devices, languages, and accessibility needs globally.

Tom: So we are moving toward an expectation that AI can function as a full-stack QA architect right out of the gate. Jane?

Jane: It raises the bar significantly for every developer who relies on these tools; they now need to think about interoperability and failure modes from day one, not as an afterthought.

Tom: This entire discussion highlights how vital comprehensive testing is for building trustworthy digital products. We covered a lot of ground today regarding "Looks Right, Works Right: A Project-Level Benchmark for Multi-Screen Mobile App Generation."

Jane: It sets a much higher standard for what we expect from the next generation of these powerful tools.

Tom: And that leaves us ready to look at how these new standards will reshape the actual development lifecycle. Next up, we'll talk about resource allocation in modern software teams.

More episodes

← Home