CivBench: Progress-Based Evaluation for LLMs' Strategic Decision-Making in Civilization V
cs.AI
Submitted: 2026-04-09
Updated: 2026-09-18
Comments: Accepted by COLM 2026
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Evaluating strategic decision-making in LLM-based agents requires generative, competitive, and longitudinal environments, yet few benchmarks provide all three, and fewer still offer evaluation
Terminology
Abstract
Evaluating strategic decision-making in LLM-based agents requires generative, competitive, and longitudinal environments, yet few benchmarks provide all three, and fewer still offer evaluation signals rich enough for long-horizon, multi-agent play. We introduce CivBench, a benchmark for LLM strategists (a model under an agentic setup) in multiplayer Civilization V. Because terminal win/loss is too sparse a signal in games spanning hundreds of turns and multiple opponents, CivBench trains estimators of a state-value function, the probability that a player eventually wins given the game state. cross 307 games with 7 LLMs and multiple CivBench agent conditions, we demonstrate CivBench's potential to estimate strategic capabilities in Civilization V as an unsaturated benchmark, reveal model-specific effects of agentic setup, and outline distinct strategic profiles not visible through outcome-only evaluation.
Sources
- When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
- Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents
- Vox Deorum: A Hybrid LLM Architecture for 4X / Grand Strategy Game AI -- Lessons from Civilization V
- GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents
- LiveAgentBench: Comprehensive Benchmarking of Agentic Systems Across 104 Real-World Challenges
- Dota 2 with Large Scale Deep Reinforcement Learning
- TMGBench: A Systematic Game Benchmark for Evaluating Strategic Reasoning Abilities of LLMs
- Digital Player: Evaluating Large Language Models based Human-like Agent in Games
- BattleAgentBench: A Benchmark for Evaluating Cooperation and Competition Capabilities of Language Models in Multi-Agent Systems
- Survey on Evaluation of LLM-based Agents
- DeepPlanning: Benchmarking Long-Horizon Agentic Planning with Verifiable Constraints
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection