SocialReasonBench: A Video-QA Benchmark for Social Reasoning with Counterfactual Narrative Videos
cs.CL, cs.CV
Submitted: 2026-08-31
Updated: 2026-08-31
Comments: Accepted to Findings of EMNLP 2026. 24 pages, 11 figures, 11 tables
Project page: https://socialreasonbench.github.io
License: http://creativecommons.org/licenses/by/4.0/
The gist: Recent advances in Large Multimodal Models (LMMs) have greatly improved video understanding, yet their ability to reason about human-centered social situations remains limited.
Terminology
Abstract
Recent advances in Large Multimodal Models (LMMs) have greatly improved video understanding, yet their ability to reason about human-centered social situations remains limited. Existing benchmarks typically rely on videos with a single observed trajectory, making it difficult to determine whether models truly understand social dynamics or merely exploit recurring narrative patterns. We introduce SocialReasonBench, a video multiple-choice QA benchmark for evaluating socially grounded reasoning in scenarios derived from interactive narratives. Built from gameplay videos of Detroit: Become Human, the benchmark leverages branching storylines where player decisions lead to alternative social outcomes that can be checked against the game's own script, flowchart, and recorded branches. We develop a multi-agent curation pipeline that localizes socially meaningful clips, grounds answer labels in game-state signals, and generates theory-guided questions with diagnostic distractors. SocialReasonBench covers seven reasoning dimensions, including intent recognition, emotional empathy, moral dilemma, counterfactual reasoning, and causal antecedent. Experiments on contemporary LMMs show that models perform reasonably well on basic social understanding but struggle with counterfactual and causal reasoning. Further ablation and diagnostic error analyses reveal that models often depend on incomplete modality cues and fall into reasoning traps such as visual shortcuts, highlighting a gap between observable event recognition and deeper reasoning over latent social states.
Sources
- ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts
- GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
- HumanVideo-MME: Benchmarking MLLMs for Human-Centric Video Understanding
- SIV-Bench: A Video Benchmark for Social Interaction Understanding and Reasoning
- Foundation Models for Video Understanding: A Survey
- Qwen3-Omni Technical Report
- Scaling Instructable Agents Across Many Simulated Worlds
- MiniCPM-V: A GPT-4V Level MLLM on Your Phone
- VideoGameBench: Can Vision-Language Models complete popular video games?
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering