GT-HarmBench: Benchmarking AI Safety Risks Through the Lens of Game Theory
cs.AI, cs.CL, cs.CY, cs.GT, cs.MA
Submitted: 2026-02-12
Updated: 2026-09-25
Code: https://github.com/causalNLP/gt-harmbench
Terminology
Sources
- Playing repeated games with Large Language Models
- AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
- When Ethics and Payoffs Diverge: LLM Agents in Morally Charged Social Dilemmas
- Media and responsible AI governance: a game-theoretic and LLM analysis
- Superintelligent Agents Pose Catastrophic Risks: Can Scientist AI Offer a Safer Path?
- Do LLMs trust AI regulation? Emerging behaviour of game-theoretic LLM agents
- FAIRGAME: a Framework for AI Agents Bias Recognition using Game Theory
- Is Power-Seeking AI an Existential Risk?
- Put Your Money Where Your Mouth Is: Evaluating Strategic Planning and Execution of LLM Agents in an Auction Arena
- DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
- From Natural Language to Extensive-Form Game Representations
- GTBench: Uncovering the Strategic Reasoning Limitations of LLMs via Game-Theoretic Evaluations
- The Llama 3 Herd of Models
- Multi-Agent Risks from Advanced AI
- Economics Arena for Large Language Models
- Corrupted by Reasoning: Reasoning Language Models Become Free-Riders in Public Goods Games
- The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning
- Can Large Language Models Trade? Testing Financial Theories with LLM Agents in Market Simulations
- Strategic Behavior of Large Language Models: Game Structure vs. Contextual Framing
- The AI risk repository: A meta-review, database, and taxonomy of risks from artificial intelligence
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection