MultivationBench: A Benchmark for Multimodal Sequential Motivation Reasoning
cs.AI
Submitted: 2026-07-29
Updated: 2026-08-29
Comments: 35 pages, 6 figures. Accepted to Findings of EMNLP 2026. Code and data: https://github.com/HKUST-KnowComp/MultivationBench
Code: https://github.com/HKUST-KnowComp/MultivationBench
License: http://creativecommons.org/licenses/by/4.0/
The gist: Multimodal Large Language Models have sparked significant interest due to their potential for social intelligence; however, their ability to perform sequential motivation reasoning remains
Terminology
Abstract
Multimodal Large Language Models have sparked significant interest due to their potential for social intelligence; however, their ability to perform sequential motivation reasoning remains insufficiently studied. Existing evaluations predominantly examine static text or isolated visual snapshots, which do not reflect the cumulative nature of real-world behavioral drivers. To address this gap, we introduce MultivationBench, a benchmark designed to rigorously evaluate multimodal motivation reasoning within story-driven visual narratives. The benchmark builds upon established psychological frameworks - Maslow's hierarchy and Reiss's basic desires - and requires models to integrate accumulated multimodal context to infer evolving motivations. Results indicate that MultivationBench presents a significant challenge: all tested models struggle to maintain consistent motivation reasoning across sequential contexts, revealing a critical disconnect between static recognition capabilities and the dynamic reasoning essential for human-like social understanding.
Sources
- Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
- Sparks of Artificial General Intelligence: Early experiments with GPT-4
- XToM: Exploring the Multilingual Theory of Mind for Large Language Models
- The Revolution of Multimodal Large Language Models: A Survey
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Text-Tuple-Table: Towards Information Integration in Text-to-Table Generation via Global Tuple Extraction
- Structuring the Unstructured: A Systematic Review of Text-to-Structure Generation for Agentic AI with a Universal Evaluation Framework
- ViSTA: Visual Storytelling using Multi-modal Adapters for Text-to-Image Diffusion Models
- Mathematical Capabilities of ChatGPT
- Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
- A Survey on LLM-as-a-Judge
- Mind the Motions: Benchmarking Theory-of-Mind in Everyday Body Language
- Privacy in Large Language Models: Attacks, Defenses and Future Directions
- Analyzing Leakage of Personally Identifiable Information in Language Models
- Towards Agentic RAG with Deep Reasoning: A Survey of RAG-Reasoning Systems in LLMs
- LLM-Hanabi: Evaluating Multi-Agent Gameplays with Theory-of-Mind and Rationale Inference in Imperfect Information Collaboration Game
- Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs
- DixitWorld: Evaluating Multimodal Abductive Reasoning in Vision-Language Models with Multi-Agent Dixit Gameplay
- SocialIQA: Commonsense Reasoning about Social Interactions
- Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection