QUACK: Questioning, Understanding, and Auditing Communicated Knowledge in Multimodal Social Deduction Agents
cs.CL, cs.AI, cs.MA
Submitted: 2026-05-26
Updated: 2026-08-31
Comments: Accepted by EMNLP 2026 Main Conference
Code: https://github.com/AAAAA-Academia-Attractions/QUACK
License: http://creativecommons.org/licenses/by/4.0/
The gist: Social deduction games have become a popular testbed for probing reasoning, deception, coordination, and belief modeling in Large Language Model (LLM) agents.
Terminology
Abstract
Social deduction games have become a popular testbed for probing reasoning, deception, coordination, and belief modeling in Large Language Model (LLM) agents. However, most environments are scored only by game outcomes such as win rates and largely remain to text-only interaction, making it difficult to tell whether an agent's language is actually grounded in what it perceived and did, or to identify the failure modes underlying its behavior. To address this gap, we introduce QUACK, an open-source environment and evaluation framework for auditing the grounding of agent language in multimodal social reasoning. QUACK evaluates agents at three levels: game outcomes, behavioral trajectories, and utterance-level consistency. Its core Statement Verification Pipeline reconstructs each agent's ground-truth trajectory from engine logs and checks every discussion claim against it, automatically flagging spatial hallucination, unsupported accusation, deception collapse, and language-action inconsistency. Evaluating three frontier VLMs in both homogeneous and cross-model adversarial settings, we find that even the strongest agent hallucinates 15.1% of its verifiable spatial claims and 11.5% of accusations are strictly unsupported. We release the full engine, evaluation framework, toolkit, and logs in https://github.com/AAAAA-Academia-Attractions/QUACK.
Sources
- Werewolf Arena: A Case Study in LLM Evaluation via Social Deduction
- LLMs are Superior Feedback Providers: Bootstrapping Reasoning for Lie Detection with Self-Generated Feedback
- Deception in LLMs: Self-Preservation and Autonomous Goals in Large Language Models
- AMONGAGENTS: Evaluating Large Language Models in the Interactive Text-Based Social Deduction Game
- The Traitors: Deception and Trust in Multi-Agent Language Model Simulations
- A Survey on Large Language Model-Based Game Agents
- Lies, Damned Lies, and Distributional Language Statistics: Persuasion and Deception with Large Language Models
- Language Agents with Reinforcement Learning for Strategic Play in the Werewolf Game
- AvalonBench: Evaluating LLMs Playing the Game of Avalon
- SocialGrid: A Benchmark for Planning and Social Reasoning in Embodied Multi-Agent Systems
- Beyond Survival: Evaluating LLMs in Social Deduction Games with Human-Aligned Strategies
- MMBench: Is Your Multi-modal Model an All-around Player?
- Avalon's Game of Thoughts: Battle Against Deception through Recursive Contemplation
- Hoodwinked: Deception and Cooperation in a Text-Based Game for Language Models
- Exploring Large Language Models for Communication Games: An Empirical Study on Werewolf
- Beyond Message Passing: A Semantic View of Agent Communication Protocols
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering