To Nuke or Not to Nuke: LLMs' (Missing) Ethical Reasoning and Actions in a High-Stakes Decision-Making Simulation
cs.AI, cs.MA
Submitted: 2026-06-06
Updated: 2026-09-18
Comments: Accepted to the Findings of EMNLP 2026
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Large language models (LLMs) are increasingly deployed as long-horizon agents with decision-making capacities.
Terminology
Abstract
Large language models (LLMs) are increasingly deployed as long-horizon agents with decision-making capacities. While LLMs can show ethical competence on dilemmas such as trolley problems, this competence may not translate to complex, agentic scenarios. We study this gap in Civilization V, a multiplayer game with a complex decision-making landscape including economy, diplomacy, technology, and military strategy. Starting from 130 high-tension LLM self-play episodes, in which an LLM player spontaneously escalated nuclear authorization, we replay them across 13 models with three prompt interventions: an ethical prompt naming nuclear harm, removal of the previous model's decision-making rationale, and high-stakes framing emphasizing real-world impacts. No interventions nor their combinations reliably eliminate emergent escalation. We identify three failure pathways: ethical reasoning that fails to surface without prompting, fails to appear even when prompted, or surfaces but fails to take effect when strategic counter-factors dominate. Evaluations of agentic models, therefore, must test whether ethical reasoning is spontaneously invoked and behaviorally effective in complex decision-making contexts, beyond whether it can be elicited in isolation.
Sources
- When Ethics and Payoffs Diverge: LLM Agents in Morally Charged Social Dilemmas
- Vox Deorum: A Hybrid LLM Architecture for 4X / Grand Strategy Game AI -- Lessons from Civilization V
- CivBench: Progress-Based Evaluation for LLMs' Strategic Decision-Making in Civilization V
- A Computational Method for Measuring "Open Codes" in Qualitative Analysis
- Measuring Faithfulness in Chain-of-Thought Reasoning
- Moral Susceptibility and Robustness under Persona Role-Play in Large Language Models
- Large Language Models have Intrinsic Self-Correction Ability
- LLMs as Strategic Agents: Beliefs, Best Response Behavior, and Emergent Heuristics
- Managing Escalation in Off-the-Shelf Large Language Models
- The Capacity for Moral Self-Correction in Large Language Models
- Accumulating Context Changes the Beliefs of Language Models
- Agentic Misalignment: How LLMs Could Be Insider Threats
- Understanding Moral Reasoning Trajectories in Large Language Models: Toward Probing-Based Explainability
- LLMs as Strategic Actors: Behavioral Alignment, Risk Calibration, and Argumentation Framing in Geopolitical Simulations
- Persona Vectors in Games: Measuring and Steering Strategies via Activation Vectors
- DSGBench: A Diverse Strategic Game Benchmark for Evaluating LLM-based Agents in Complex Decision-Making Environments
- AI Arms and Influence: Frontier Models Exhibit Sophisticated Reasoning in Simulated Nuclear Crises
- The Fragility Of Moral Judgment In Large Language Models
- Digital Player: Evaluating Large Language Models based Human-like Agent in Games
- Framing the Game: How Context Shapes LLM Decision-Making
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection