CivBench: Progress-Based Evaluation for LLMs' Strategic Decision-Making in Civilization V

arXiv:2604.07733 · cs.AI · Submitted 2026-04-09 · Read on arXiv

cs.AI

Submitted: 2026-04-09

Updated: 2026-09-18

Comments: Accepted by COLM 2026

License: http://creativecommons.org/licenses/by-nc-sa/4.0/

The gist: Evaluating strategic decision-making in LLM-based agents requires generative, competitive, and longitudinal environments, yet few benchmarks provide all three, and fewer still offer evaluation

Terminology

Abstract

Evaluating strategic decision-making in LLM-based agents requires generative, competitive, and longitudinal environments, yet few benchmarks provide all three, and fewer still offer evaluation signals rich enough for long-horizon, multi-agent play. We introduce CivBench, a benchmark for LLM strategists (a model under an agentic setup) in multiplayer Civilization V. Because terminal win/loss is too sparse a signal in games spanning hundreds of turns and multiple opponents, CivBench trains estimators of a state-value function, the probability that a player eventually wins given the game state. cross 307 games with 7 LLMs and multiple CivBench agent conditions, we demonstrate CivBench's potential to estimate strategic capabilities in Civilization V as an unsaturated benchmark, reveal model-specific effects of agentic setup, and outline distinct strategic profiles not visible through outcome-only evaluation.

Sources

Related papers