SocialMaze: A Benchmark for Evaluating and Enhancing Social Reasoning in Large Language Models in Complex Social Environments
cs.CL
Submitted: 2025-05-29
Updated: 2026-09-20
Terminology
Sources
- ShieldGemma: Generative AI Content Moderation Based on Gemma
- Hate Personified: Investigating the role of LLMs in content moderation
- LLM-Assisted Content Analysis: Using Large Language Models to Support Deductive Coding
- The Stepwise Deception: Simulating the Evolution from True News to Fake News with LLM Agents
- Theory of Mind for Multi-Agent Collaboration via Large Language Models
- Exploring Large Language Models for Word Games:Who is the Spy?
- On the Trustworthiness of Generative Foundation Models: Guideline, Assessment, and Perspective
- MJ-Bench: Is Your Multimodal Reward Model Really a Good Judge for Text-to-Image Generation?
- Injecting Domain-Specific Knowledge into Large Language Models: A Comprehensive Survey
- Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs
- SocialIQA: Commonsense Reasoning about Social Interactions
- CREAK: A Dataset for Commonsense Reasoning over Entity Knowledge
- GoEmotions: A Dataset of Fine-Grained Emotions
- Clever Hans or Neural Theory of Mind? Stress Testing Social Reasoning in Large Language Models
- Evaluating Theory of Mind in Question Answering
- Social Chemistry 101: Learning to Reason about Social and Moral Norms
- DebateBench: A Challenging Long Context Reasoning Benchmark For Large Language Models
- Mastering the Game of No-Press Diplomacy via Human-Regularized Reinforcement Learning and Planning
- AvalonBench: Evaluating LLMs Playing the Game of Avalon
- Werewolf Arena: A Case Study in LLM Evaluation via Social Deduction
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering