GUI Agents for Continual Game Generation

arXiv:2605.28258 · cs.SE, cs.AI, cs.CV, cs.HC · Submitted 2026-05-27 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "GUI Agents for Continual Game Generation".

Jane: Generating a game is not the same as making one that can be played,

Tom: First, who's behind it and why it matters.

Paper summary: Tom: Welcome back to the show, everyone! We've got a fascinating paper today that really digs into how we actually test the games AI makes. It’s titled "GUI Agents for Continual Game Generation," and it’s something we need to talk about because it shifts how we think about creating interactive code.

Jane: I agree, Tom; this paper tackles the core problem that making a game isn't the same as being able to play one. It points out that current methods often treat game generation as just a one-shot translation from a prompt to an artifact, which means they miss those crucial interaction failures that happen during actual gameplay.

Lu: Exactly, and what this work claims is that evaluating and improving game generation absolutely requires a player involved in the process. They propose two distinct roles for graphical user interface agents to handle this need for playtesting: one as an objective evaluator through something called PlaytestArena, and another as a subjective playtester via Play2Code.

Meng: So they’re suggesting that we need an agent that can actually load the generated game into a browser and judge it against some expected behaviors, rather than just looking at the code itself. That moves us toward testing in the actual environment.

Lalam: And what this means for culture is huge; if we can build systems where AI generates something that is objectively tested by an agent, it opens up new ways for our creative output to be refined iteratively rather than through a single guess.

Tom: Right, and the paper argues that these two roles—the objective evaluator and the subjective playtester—are key to improving how we generate interactive code. They set up PlaytestArena as this environment pairing two hundred game generation tasks across eight genres with rubrics defining what you expect to see during play <ref:2605.28258#pg0,game generation tasks across eight genres with rubrics>.

Jane: That objective evaluation setup sounds really practical, Tom; it’s like having a standardized way to measure success before you even get into the deeper refinement process. It helps study how generation methods perform when the focus moves from just looking at the code to actually playing with it.

Lu: And then they propose Play2Code as the subjective playtester, which is where things get really interesting because it bridges game generation and playtesting by running a game agent and a GUI agent in a continuous loop with shared memory. This turns what used to be just generation into an ongoing dialogue between coding and playing.

Meng: A sustained loop sounds intensive from an engineering standpoint; how does that shared memory system actually manage the flow of information between the code writing part and the play testing part without causing bottlenecks?

Lalam: From my side, I see this as a way to build a sort of cultural feedback mechanism where every iteration is shaped by what played before it, which really pushes our generative capabilities forward.

Paper summary: Tom: That’s a great question about the engineering flow, Meng; the paper describes several interconnected components designed for this continual refinement process. We have a Game Agent that writes, debugs, and patches game code through five phases including design and verification.

Jane: And then paired with that is the GUI Agent which loads the runtime in a browser without access to the source code, relying only on a guide and memory to conduct its initial observation and interactive playtesting.

Lu: The system relies on a three-layer memory structure—Episode Memory for in-task experience, Skill Memory for cross-task isolation, and World Memory for shared experience across tasks. This allows the Game Agent to query this memory at the start of each task based on its archetype and role.

Meng: So the architecture is designed to accumulate experience across tasks while keeping certain types of knowledge isolated or shared depending on what's needed for debugging or refinement? That sounds like a complex state management problem that needs careful design.

Lalam: It’s about creating an evolving system where the agents don't just start fresh every time; they build upon what they’ve already seen, which is really powerful for developing more robust systems over time.

Tom: And the results from their experiments on Play2Code across different backbones like GPT-five point four, Sonnet four point six, and Kimi K2 point 5 show some significant performance gains when compared to single-pass methods in other approaches they tested.

Jane: Those performance metrics are encouraging; achieving a "sixty-six point eight percent rubric pass-rate" is a solid indicator of how much better this continual evolution approach is at passing the defined criteria than simpler generation methods. They showed substantial improvements over baselines, specifically gaining thirty-seven point one and fourteen point six points respectively compared to those prior methods cited in Hu et al., twenty <ref:2605.28258#pg1>.

Lu: What they found in their analysis was that the feedback from the GUI playtester is more traceable than a human report, but it still has idiosyncrasies that remind us of how real human testers operate. This suggests we get structured data even when the tester isn't human.

Meng: That traceability is valuable for debugging, I can see that; having actionable fix lists generated from the play testing observation rather than just a vague failure report makes it much more useful for an engineer to follow.

Lalam: The fact that they confirmed the GUI Agent's role is essential—that removing it caused a substantial drop in rubric score—really shows us that execution-based feedback is what surfaces those runtime failures and interaction-layer bugs we struggle with.

Tom: It really does, and when we looked at how they compared this GUI-driven refinement against human playtesting on fifty moderate-complexity tasks, the results showed that human refinement still achieved higher rubric scores across all rounds.

Jane: That comparison is important because it tells us that while the GUI Agent provides a useful signal, it hasn't quite reached the level of subjective understanding that a human player brings to things. They noted this limitation when discussing how complex games are involved in their testing.

Paper summary: Lu: The paper points out that the effectiveness of Play2Code is also modulated by game complexity; low-complexity games tend to have simple, self-contained mechanics, but high-complexity games introduce intricate conditional logic or emergent interactions that are hard for the GUI agent to trigger reliably.

Meng: That makes sense from a practical standpoint; if the game state transitions become too complicated or unpredictable, an automated agent might struggle to reliably navigate those complex pathways and provide useful feedback.

Lalam: It suggests that the GUI agents carry something like a taste at the level of what they notice, even if it's not perfect human intuition yet; they are picking up on patterns in state transitions.

Tom: So, while Play2Code is a strong signal for refinement, it doesn't replace human testers entirely because it can’t capture the feeling of difficulty or surprise that comes with truly complex interactions.

Jane: That distinction is vital; they admitted that GUI agents fall short when trying to deliver that "felt sense of difficulty, frustration, boredom, and surprise" that a person experiences while playing. It’s a limitation they have openly stated about their current system.

Lu: The divergence in feedback categories across different backbones used in their experiments also suggests that different AI models notice different aspects of the game experience; it's not just one way to see what's wrong.

Meng: That tells us we might need a more diverse set of evaluation tools, perhaps combining structured agent feedback with qualitative human input to get the full picture.

Lalam: If we combine that structured feedback from Play2Code with the deep, nuanced understanding from human playtesting, I think we can build a system that's incredibly effective at refining these kinds of generative tasks.

Tom: So, to wrap up this look at "GUI Agents for Continual Game Generation," the authors are showing us how to move beyond single-pass generation by creating a dialogue between coding and playing through systems like PlaytestArena and Play2Code.

Jane: They’re essentially proposing a framework where we use agents that can both objectively test outputs and subjectively play them to guide the next version of the code, which is a really interesting way to approach iterative development.

Lu: The implication for the future is that we start treating game generation not as an end-to-end synthesis problem, but as a continual evolution process where player interaction is baked into the design cycle.

Meng: From my side, I see the practical impact in reducing the time spent debugging those tricky interaction bugs early on because this system forces that testing to happen constantly during development.

Lalam: For our culture here at the startup, this means we can generate things faster and with higher quality assurance built right into the generation pipeline, making our output more reliable and sophisticated.

Tom: That’s a great summary of the core idea; it’s about embedding testing directly into the creation loop instead of just checking a finished product once.

Jane: And that ties nicely back to how we can use these agents to systematically improve our code generation capabilities for interactive experiences, showing us a path forward.

Conclusion: Tom: So, to wrap up our discussion on "GUI Agents for Continual Game Generation," we've seen how these agents create a feedback loop between generating code and actually playing with it across different systems.

Jane: It really boils down to this idea of treating game creation as an ongoing conversation rather than just a single attempt at writing everything at once.

Lu: The authors are proposing these GUI agents—one for objective evaluation and one for subjective playtesting—to turn those gameplay observations into actionable refinements for the AI.

Meng: From my side, I’m focused on how this structure handles the actual engineering pipeline, making sure the feedback we get is something we can reliably use to patch code.

Lalam: The vision here is that by embedding playtesting directly into the generation cycle, we can build systems that evolve with every iteration based on real interaction data.

Tom: It's fascinating how they've structured it, moving from a static evaluation setup like PlaytestArena to a dynamic feedback loop called Play2Code.

Jane: They’re essentially showing us that for interactive code generation, we need agents that can both judge what's wrong objectively and play things subjectively.

Lu: The authors are really pushing the idea that this structured, traceable feedback from the GUI agent is a scalable way to improve how we handle complex interactive systems.

Meng: It makes sense that they’d focus on traceability; if we can pinpoint exactly where the game failed during play, it cuts down on endless debugging cycles.

Lalam: This kind of iterative refinement could fundamentally alter how we approach building sophisticated interactive experiences in the future, moving us away from just hoping a single generation produces a playable result.

Tom: And that leads right into what this means for our entire creative process; we're talking about building quality assurance directly into the creation pipeline itself.

Yixu Huang, Bo Li, Na Li, Zhe Wang, Kaijie Chen, Haonan Ge, Qingyi Si

Fudan University

cs.SE, cs.AI, cs.CV, cs.HC

Submitted: 2026-05-27

Updated: 2026-10-02

Project page: https://continual-game-generation.vercel.app

Importance score: 88/100

The gist: Generating a game is not the same as making one that can be played, and this work investigates how graphical user interface (GUI) agents can serve as objective evaluators and subjective playtesters

Key concepts

PlaytestArena
This environment pairs 200 different game generation tasks across eight genres with specific rules (rubrics) defining what good play looks like. A GUI agent loads the generated game, plays it by clicking and using keys, and judges the output against these rules. It helps study how code performance changes when evaluation shifts from just looking at the code to actually playing the finished game.
Play2Code
This system creates a continuous loop between a Game Agent (which writes/patches code) and a GUI Agent (which plays the game). They share memory, allowing playtesting observations to directly shape the next version of the game. This treats generation as an ongoing dialogue between coding and playing, where each build is improved based on how well it was played.
GUI Agent
This agent loads a generated game into a web browser and plays through it using clicks and keys without needing access to the source code. It relies on a guide and memory to interact with the game. Its role is crucial because its execution-based feedback surfaces bugs related to how the game actually runs, which is essential for refining code.
Memory System
The system uses three layers of memory: Episode Memory for current tasks, Skill Memory for cross-task learning between agents, and World Memory shared across all tasks. This structure allows the Game Agent to query relevant experience before starting a new task, helping it accumulate knowledge and refine its generation process over time.

Terminology

Summary

Generating a game is not the same as making one that can be played, and this work investigates how graphical user interface (GUI) agents can serve as objective evaluators and subjective playtesters to improve interactive code generation. The central argument is that evaluating and refining game generation requires a player, leading to the proposal of two distinct roles for GUI agents: an objective evaluator via PlaytestArena and a subjective playtester via Play2Code.

Objective Evaluation with PlaytestArena

The first role explored is that of an objective evaluator, introduced through PlaytestArena. This environment pairs 200 browser-based game generation tasks across eight genres with rubrics defining expected in-play behaviors. A GUI agent serves as the adjudicator by loading each generated build into a browser, playing it through clicks and keys, and judging each rubric criterion based on what it observes during play. This setup is useful for studying how generation methods perform when the unit of evaluation shifts from code to play.

Subjective Playtesting with Play2Code

The second role proposes Play2Code, a system that bridges game generation and playtesting by operating a game agent and a GUI agent in a sustained loop with shared memory. This turns game generation into a dialogue between coding and playing, treating it as continual evolution where each build is shaped by the previous one's play experience. The agents operate around a shared runtime, with the GUI agent surfacing observations, and the game agent using them to refine the next iteration.

Key System Components

The Play2Code system involves several interconnected components designed for continual refinement:

  1. A Game Agent that writes, debugs, and patches game code, operating in a multi-phase workflow including game design, asset generation, code implementation, build-time verification, and memory capture. It uses tools like file-system access to manage the artifact.

  2. A GUI Agent that loads the runtime in a browser and plays through it without source code access, relying only on a game guide and memory. Its workflow includes initial observation, game start, interactive playtesting, assessment, culminating in an output that is parsed into a play summary and actionable fix list.

  3. A three-layer memory system: Episode Memory (in-task), Skill Memory (cross-task/agent-isolated), and World Memory (cross-task/shared). These layers allow the system to accumulate experience, with the Game Agent querying memory at the start of each task scoped by archetype and role.

Experimental Results and Findings

Experiments on Play2Code across three frontier backbones (GPT-5.4, Sonnet 4.6, and Kimi K2.5) show significant performance gains over single-pass methods. Play2Code achieves a 66.8% rubric pass-rate, improving over baselines by substantial margins (37.1 and 14.6 points). Further analysis reveals that GUI playtester feedback is more traceable than a human report, yet idiosyncratic in ways reminiscent of human testers themselves. Ablation studies confirm that the GUI Agent’s role is essential, as removing it causes a substantial drop in rubric score, underscoring that grounded, execution-based feedback surfaces runtime failures and interaction-layer bugs.

Human Comparison and Complexity

The work compares GUI-driven refinement against human playtesting on 50 Moderate-complexity tasks. Results show that human-driven refinement achieves higher rubric scores than GUI agent-driven refinement across all rounds, confirming the GUI Agent has not yet closed the gap. The effectiveness of Play2Code is also modulated by game complexity: Low-complexity games tend to involve simple, selfcontained mechanics, while High-complexity games involve intricate conditional logic, multiphase state transitions, or emergent interactions that are difficult for the GUI agent to trigger reliably. This suggests that GUI agents carry something like a taste at the level of what they notice.

Conclusion on Playtesting

The paper concludes that playtesting provides a usable optimization signal for code generation, turning gameplay observations into an iterative refinement process. While GUI agents are not yet perfect substitutes for human playtesters—as they fall short in delivering the felt sense of difficulty, frustration, boredom, and surprise—their ability to produce structured, traceable feedback makes them a scalable and competitive refinement signal. The divergence in feedback categories across different backbones suggests that different agents notice different aspects of the game experience.

The gist

Play2Code achieves a 66.8% rubric pass-rate by coupling a game agent with a GUI agent in a sustained loop, turning generation into a dialogue between coding and playing.

How it works

  1. A Game Agent writes, debugs, and patches game code, operating in five phases: game design, asset generation, code implementation, build-time verification, and memory capture.

Improvements for AI systems

Here are specific improvements for AI systems, derived from the principles and findings of the provided research paper, and what those improved systems can achieve:


The core improvement is shifting from one-shot generation to a continual evolution paradigm driven by interactive playtesting. This involves integrating a dual-agent system (Game Agent + GUI Agent) operating in a shared memory loop.

Here are the specific improvements and capabilities for the resulting AI systems:

  1. Replaces static code generation with an iterative, feedback-driven development loop:

  2. A Game Agent that writes, debugs, and patches game code based on previous builds and external feedback (play summaries).

  3. A GUI Agent that acts as a subjective playtester, loading the build in a browser to observe rendered gameplay via clicks and keys, reasoning about the interface state, and producing structured feedback (summary + actionable fix list) grounded in observed behavior rather than static code inspection.

  4. A Shared Memory System that accumulates experience across rounds and tasks, enabling both agents to learn:

  5. Skill Memory (agent-specific patterns/strategies for code generation).

  6. World Memory (cross-task design principles and common archetypes of game mechanics).

The improved AI system (Play2Code) can achieve the following specific capabilities:

  1. Generate functional, playable artifacts that satisfy complex, multi-dimensional rubrics defined by human experts across various game genres (e.g., puzzle, strategy, card games).

  2. Improve artifact quality monotonically over iterative refinement rounds (e.g., achieving a 66.8% rubric pass-rate over multiple iterations) because the AI is guided by empirical play rather than just static code checks.

  3. Identify and correct subtle, interaction-level failures that are invisible to traditional code inspection or compilation tests (such as unresponsive controls, incorrect state transitions, or timing errors in platformers).

  4. Develop genre-specific heuristics and patterns (via Skill Memory) to accelerate development on recurring implementation pitfalls within a specific game type.

  5. Transfer successful design knowledge from one generation task to another via World Memory, allowing the system to leverage learned archetypes and universal design principles for new games.

  6. Bridge the gap between AI code-generation and human playtesting by producing feedback that is traceable (anchored in observable frames) yet idiosyncratic (reflecting nuanced game feel or aesthetic judgments).

  7. Optimize development efficiency by allowing the Game Agent to intelligently prioritize which suggested fixes from the GUI Agent's report are relevant and feasible for implementation in the next round.

Sources

Related papers