A2Z GameSpec-Bench: How Faithfully Can Coding Agents Generate Games from Game Design Specifications?
cs.AI, cs.CV
Submitted: 2026-09-30
Updated: 2026-09-30
Project page: https://a2z-gamespec-bench.github.io
Terminology
Sources
- WebGrader: Training LLMs for Web Development with Self-Evolving Programmatic Grader
- GameDevBench: Evaluating Agentic Capabilities Through Game Development
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
- WebVR: Benchmarking Multimodal LLMs for WebPage Recreation from Videos via Human-Aligned Visual Rubrics
- DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
- MUSE: Benchmarking Manufacturable, Functional, and Assemblable Text-to-CAD Generation
- Automatic Detection of Causality in Requirement Artifacts: the CiRA Approach
- GLM-5: from Vibe Coding to Agentic Engineering
- GUI Agents for Continual Game Generation
- GameGen-Verifier: Parallel Keypoint-Based Verification for LLM-Generated Games via Runtime State Injection
- OpenGame: Open Agentic Coding for Games
- Kimi K2: Open Agentic Intelligence
- GameEngineBench: Evaluating Coding Agents on Real C++ Runtime Environments
- Meta-Harness: End-to-End Optimization of Model Harnesses
- OmniGameArena: A Unified UE5 Benchmark for VLM Game Agents with Improvement Dynamics
- GameCraft-Bench: Can Agents Build Playable Games End-to-End in a Real Game Engine?
- WebRISE: Requirement-Induced State Evaluation for MLLM-Generated Web Artifacts
- GameWorld: Towards Standardized and Verifiable Evaluation of Multimodal Game Agents
- LiveEvalBench: Toward Open-World Evaluation for Web Generation
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection