SocialOmni: Benchmarking Audio-Visual Social Interactivity in Omni Models
cs.AI
Submitted: 2026-03-17
Updated: 2026-09-20
Comments: 23 pages, including appendix. Updated evaluation protocol and results; reproducibility materials included as ancillary files. Code: https://github.com/MAC-AutoML/SocialOmni . Dataset: https://huggingface.co/datasets/alexisty/SocialOmni
Code: https://github.com/MAC-AutoML/SocialOmni
License: http://creativecommons.org/licenses/by/4.0/
The gist: Evaluating omni-modal large language models (OLMs) in multi-party dialogue requires more than answer correctness on pre-segmented inputs.
Terminology
Abstract
Evaluating omni-modal large language models (OLMs) in multi-party dialogue requires more than answer correctness on pre-segmented inputs. We introduce SocialOmni, an offline diagnostic benchmark that separates three turn-level decisions: identifying who is speaking, deciding when a designated participant should enter at an annotated query time, and determining how that participant should continue the dialogue. SocialOmni contains 2,000 perception items and a quality-controlled core split of 200 interaction-generation items, including naturally occurring speaker-visibility mismatches. Each item is independently checked by three human annotators. Evaluated systems receive only query-time-bounded multimodal evidence, while manually verified reference continuations are reserved for response judging. A complete three-judge ensemble scores every eligible response, with leave-one-judge-out and family-sensitivity audits. Across 11 OLMs, rankings vary substantially by axis, and coverage-adjusted scores reveal when high conditional response quality depends on selective turn entry. The protocol does not measure persistent streaming state or wall-clock latency.
Sources
- Flamingo: a Visual Language Model for Few-Shot Learning
- Talking Turns: Benchmarking Audio Foundation Models on Turn-Taking Dynamics
- JointAVBench: A Benchmark for Joint Audio-Visual Reasoning Evaluation
- Event-Anchored Frame Selection for Effective Long-Video Understanding
- Wavelet-based Frame Selection by Detecting Semantic Boundary for Long Video Understanding
- AMUSE: Audio-Visual Benchmark and Alignment Framework for Agentic Multi-Speaker Understanding
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction
- FunASR: A Fundamental End-to-End Speech Recognition Toolkit
- WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs
- Determining Layer-wise Sparsity for Large Language Models Through a Theoretical Perspective
- Dynamic Low-Rank Sparse Adaptation for Large Language Models
- GPT-4o System Card
- SIV-Bench: A Video Benchmark for Social Interaction Understanding and Reasoning
- SEED-Bench-2: Benchmarking Multimodal Large Language Models
- OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs
- Baichuan-Omni-1.5 Technical Report
- OmniBench: Towards The Future of Universal Omni-Language Models
- Full-Duplex-Bench: A Benchmark to Evaluate Full-duplex Spoken Dialogue Models on Turn-taking Capabilities
- WearVox: An Egocentric Multichannel Voice Assistant Benchmark for Wearables
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection