MultivationBench: A Benchmark for Multimodal Sequential Motivation Reasoning

arXiv:2607.26465 · cs.AI · Submitted 2026-07-29 · Read on arXiv

cs.AI

Submitted: 2026-07-29

Updated: 2026-08-29

Comments: 35 pages, 6 figures. Accepted to Findings of EMNLP 2026. Code and data: https://github.com/HKUST-KnowComp/MultivationBench

Code: https://github.com/HKUST-KnowComp/MultivationBench

License: http://creativecommons.org/licenses/by/4.0/

The gist: Multimodal Large Language Models have sparked significant interest due to their potential for social intelligence; however, their ability to perform sequential motivation reasoning remains

Terminology

Abstract

Multimodal Large Language Models have sparked significant interest due to their potential for social intelligence; however, their ability to perform sequential motivation reasoning remains insufficiently studied. Existing evaluations predominantly examine static text or isolated visual snapshots, which do not reflect the cumulative nature of real-world behavioral drivers. To address this gap, we introduce MultivationBench, a benchmark designed to rigorously evaluate multimodal motivation reasoning within story-driven visual narratives. The benchmark builds upon established psychological frameworks - Maslow's hierarchy and Reiss's basic desires - and requires models to integrate accumulated multimodal context to infer evolving motivations. Results indicate that MultivationBench presents a significant challenge: all tested models struggle to maintain consistent motivation reasoning across sequential contexts, revealing a critical disconnect between static recognition capabilities and the dynamic reasoning essential for human-like social understanding.

Sources

Related papers