Can AI Agents Learn Their Way to the Top? Evaluating Heuristic Learning in a Long-Running Game Agent Competition
cs.AI, cs.CL
Submitted: 2026-10-08
Updated: 2026-10-08
Code: https://github.com/THU-CST-SAST/AAArena
Project page: https://aaarena.net
Terminology
Sources
- GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents
- Can Vibe Coding Beat Graduate CS Students? An LLM vs. Human Coding Tournament on Market-driven Strategic Planning
- Game Arena: Strategic LLM Evaluation in Competitive Environments
- From Code to Play: Benchmarking Program Search for Games Using Large Language Models
- CATArena: Evaluating Evolutionary Capabilities of Code Agents via Iterative Tournaments
- Heuristic Learning for Active Flow Control Using Coding Agents
- TextArena
- lmgame-Bench: How Good are LLMs at Playing Games?
- GameArena: Evaluating LLM Reasoning through Live Computer Games
- Automated Design of Agentic Systems
- ZeroSumEval: Scaling LLM Evaluation with Inter-Model Competition
- Learning Game-Playing Agents with Generative Code Optimization
- Eureka: Human-Level Reward Design via Coding Large Language Models
- AlphaEvolve: A coding agent for scientific and algorithmic discovery
- BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games
- ProxyWar: Dynamic Assessment of LLM Code Generation in Game Arenas
- EvoPolicyGym: Evaluating Autonomous Policy Evolution in Interactive Environments
- OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
- SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection