Compiled Agency: Coding Agents as Game AI Researchers -- from a Roguelike to StarCraft II and Civilization
cs.AI, cs.CL
Submitted: 2026-07-17
Updated: 2026-09-26
License: http://creativecommons.org/licenses/by/4.0/
The gist: LLM agents have repeatedly struggled to convert knowledge of a game into competent play, even when researchers build the agent around the model - supplying perception, memory, skill libraries,
Terminology
Abstract
LLM agents have repeatedly struggled to convert knowledge of a game into competent play, even when researchers build the agent around the model - supplying perception, memory, skill libraries, planners, or executable-policy scaffolds. Rapid progress in coding agents raises two sharper questions: can frontier models now win games at all, and can they win them unaided, building the entire player themselves? We introduce Gauntlet, a develop-freeze-evaluate framework that ports games from small arcades to full commercial-scale titles, behind one deliberately bare contract: a general-purpose coding agent receives a game description, a raw observation/action interface, and an empty policy file - no strategy, no algorithm, no architecture. In a single autonomous session the agent experiments with the live game and engineers a standalone controller; we freeze the result and score it on held-out instances with zero model calls during play. On an unpublished procedural roguelike, held-out success spans 0-86 percent and exposes a sharp generational threshold: every observed session of a newest-generation system outperforms the best session of its predecessor. At full-game scale, a compiled raw-API controller defeats every fair StarCraft II built-in AI and two cheating variants, and single-session programs win complete games of Civilization (Freeciv) by total conquest on held-out seeds. Though at modest rates against novice AI, this is a first: no prior language-agent system had won full games of this genre standalone, without per-turn model calls and a hand-crafted tactical layer. Frontier coding agents begin to track long-horizon strategy. The frozen programs are inspectable. We call this capability compiled agency: development experience compiled into a persistent executable agent whose architecture is built by the model.
Sources
- On Randomness in Agentic Evals
- MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
- Vox Deorum: A Hybrid LLM Architecture for 4X / Grand Strategy Game AI -- Lessons from Civilization V
- lmgame-Bench: How Good are LLMs at Playing Games?
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
- Code World Models for General Game Playing
- GVGAI-LLM: Evaluating Large Language Model Agents with Infinite Games
- Hierarchical Expert Prompt for Large-Language-Model: An Approach Defeat Elite AI in TextStarCraft II for the First Time
- PolicyEvolve: Evolving Programmatic Policies by LLMs for multi-player games via Population-Based Training
- Large Language Models Play StarCraft II: Benchmarks and A Chain of Summarization Approach
- CivRealm: A Learning and Reasoning Odyssey in Civilization for Decision-Making Agents
- Frontier Coding Agents Can Now Implement an AlphaZero Self-Play Machine Learning Pipeline For Connect Four That Performs Comparably to an External Solver
- TStarBots: Defeating the Cheating Level Builtin AI in StarCraft II in the Full Game
- Cradle: Empowering Foundation Agents Towards General Computer Control
- Atari-GPT: Benchmarking Multimodal Large Language Models as Low-Level Policies in Atari Games
- Text2Reward: Reward Shaping with Language Models for Reinforcement Learning
- AI Gamestore: Scalable, Open-Ended Evaluation of Machine General Intelligence with Human Games
- EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection