SocialOmni: Benchmarking Audio-Visual Social Interactivity in Omni Models

arXiv:2603.16859 · cs.AI · Submitted 2026-03-17 · Read on arXiv

cs.AI

Submitted: 2026-03-17

Updated: 2026-09-20

Comments: 23 pages, including appendix. Updated evaluation protocol and results; reproducibility materials included as ancillary files. Code: https://github.com/MAC-AutoML/SocialOmni . Dataset: https://huggingface.co/datasets/alexisty/SocialOmni

Code: https://github.com/MAC-AutoML/SocialOmni

License: http://creativecommons.org/licenses/by/4.0/

The gist: Evaluating omni-modal large language models (OLMs) in multi-party dialogue requires more than answer correctness on pre-segmented inputs.

Terminology

Abstract

Evaluating omni-modal large language models (OLMs) in multi-party dialogue requires more than answer correctness on pre-segmented inputs. We introduce SocialOmni, an offline diagnostic benchmark that separates three turn-level decisions: identifying who is speaking, deciding when a designated participant should enter at an annotated query time, and determining how that participant should continue the dialogue. SocialOmni contains 2,000 perception items and a quality-controlled core split of 200 interaction-generation items, including naturally occurring speaker-visibility mismatches. Each item is independently checked by three human annotators. Evaluated systems receive only query-time-bounded multimodal evidence, while manually verified reference continuations are reserved for response judging. A complete three-judge ensemble scores every eligible response, with leave-one-judge-out and family-sensitivity audits. Across 11 OLMs, rankings vary substantially by axis, and coverage-adjusted scores reveal when high conditional response quality depends on selective turn entry. The protocol does not measure persistent streaming state or wall-clock latency.

Sources

Related papers