MCBench: A Multicontext Safety Assessment Benchmark for Omni Large Language Models

arXiv:2606.05177 · cs.CL, cs.AI, eess.AS · Submitted 2026-08-17 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio.

Tom: Next we'll be talking about the paper "MCBench: A Multicontext Safety Assessment Benchmark for Omni Large Language Models".

Jane: The paper was written by Manh Luong, Tamas Abraham, Junae Kim, Amar Kaur, Rollin Omari et al. from Monash University and Defence Science and Technology Group.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Summary: Tom: The summary section highlights some really critical failures in the current state-of-the-art LLMs. It's not that the AI can't see or hear us; it’s that it can’t reason across modalities.

Jane: The authors found that while these omni models perform well on clear physical danger, like a fire or a car crash, they struggle with risks that are more abstract.

Lu: That is the failure to integrate the cues; for example, they do better when there's a visible hazard than when there is social or legal liability.

Meng: It seems like the visual and acoustic cues provide concrete information that helps the models, but subtle issues require a kind of "multicontext" thinking that current systems lack.

Lalam: The implication here for our culture is that we are not yet ready to deploy AI in sensitive areas requiring complex societal judgment because it cannot handle nuance.

Tom: The research traces show the AI can extract individual information—it knows what's happening in the picture and what was said—but it often fails to combine those cues into a single accurate safety prediction.

Jane: Think of it like this: the model sees a person, hears them say something concerning, and reads text on a screen, but it doesn't know how to weigh those three distinct pieces of information against each other.

Lu: That lack of effective integration is the core issue; they are failing to build a cohesive picture from fragments.

Meng: This means we need architectures that can synthesize data streams rather than just running parallel processing streams for each modality.

Lalam: If the AI cannot handle those subtle, non-physical risks, it will fail spectacularly when dealing with complex human interaction and social dynamics.

Tom: So, the paper is saying that current omni LLMs lack a robust way to reason across different modalities in safety-critical settings.

Jane: It's a wake-up call for us realizing that we aren't just building bigger models; we are building more interconnected systems.

Lu: We are discovering fundamental limitations in the design of how these multimodal inputs are processed and weighted.

Meng: The engineering challenge is clear: we need to find a way to force balanced multicontext integration before this technology can be trusted widely.

Improvements: Tom: The findings suggest that if we improve how the models integrate information, they will become much more reliable in these subtle scenarios.

Jane: The authors highlight that when a model is given clear cues, its performance improves drastically, but it also shows problems with oversensitivity when it's not sure what to do.

Lu: This oversensitivity—where the model focuses too hard on one small signal and ignores all the surrounding context—is a huge area for improvement.

Meng: In practice, this means we need training strategies that prevent "isolated cues" from driving the entire safety judgment without considering other modalities at play.

Lalam: We need AI that is not just reactive to what it hears or sees, but proactive in how it weighs all available evidence towards a balanced conclusion.

Tom: The paper suggests that we should build systems where the model is forced to look at the entire context before making a decision, not just focusing on one alarming piece of data.

Jane: It's about teaching the AI to be more holistic and less prone to jumping to conclusions based on a single trigger.

Lu: We need mechanisms that allow us to see how one modality interacts with others—how does the sound influence the visual interpretation?

Meng: From an engineering view, this means designing training objectives that reward balanced multicontextual reasoning over simply achieving high accuracy on individual input types.

Lalam: If we can teach it to weigh every piece of information equally, the impact on public trust and safety applications will be enormous.

Tom: So, we’ are talking about moving beyond just teaching us to see or hear things correctly; we’re talking about making the sure how they *think* about* what they see and hear.

Jane: A shift from merely improving perception to fundamentally improving the reasoning process itself is a major conceptual leap.

The Impact: Tom: This benchmark, MCBench: A Multicontext Safety Assessment Benchmark for Omni Large Language Models, is designed to be a rigorous testbed for safety advancement.

Jane: It provides a standard that forces us to evaluate how reliable these powerful AI systems are when we need them most in real-world complexity.

Lu: The theoretical implication is that the current architecture of LLMs lacks the necessary robust cross-modal reasoning capabilities to handle truly complex, multi-layered scenarios.

Meng: Practical implications are huge; if we can’t trust the AI with subtle safety judgments, we can't deploy it in high-stakes industries like healthcare or autonomous driving.

Lalam: We are seeing how a foundational change in how we test our AI will profoundly affect how humans interact with technology and what kind of decisions they trust.

Tom: The fact that the model struggles with social and legal harm but succeeds with property damage is telling us about our current biases in testing.

Jane: It seems like the models are better at physical threats because those cues are very concrete, but bad at judging human intent or complex legal situations.

Lu: This suggests a gap between our ability finding visual cues and our ability understanding human systems and social structures.

Meng: We need to build AI that is as good at understanding a courtroom scenario as it is at recognizing a fire alarm, which is quite different.

Lalam: The goal isn' to ensure the AI doesn't just mimic human behavior, but genuinely understands the context of ethical and legal consequences.

Tom: It’s forcing us to build safety into every single component rather than trying to patch it on later in a multimodal model.

Jane: This is about creating an entirely new level of reliability that will benefit everyone who relies on these powerful AI systems.

Conclusion: Tom: So, to summarize, we have a powerful new tool that provides a rigorous testbed for advancing safe AI development.

Jane: We learned that current models struggle with nuanced safety risks because they can't integrate cues across modalities effectively.

Lu: The core of the finding is the lack of robust multicontextual reasoning, which is a fundamental weakness in current AI architecture.

Meng: It also showed that we are oversensitive to isolated cues even when we have ground-truth information, which is a major flaw in operational design.

Lalam: This benchmark pushes us toward an AI that truly understands context, not just one single element of it.

Tom: We really appreciate the authors for creating this new standard and providing such detailed findings on the challenges facing our current LLMs.

Jane: It’s a necessary step toward ensuring we are building reliable, trustworthy systems for a future where AI is ubiquitous.

Manh Luong, Tamas Abraham, Junae Kim, Amar Kaur, Rollin Omari, Gholamreza Haffari, Trang Vu, Lizhen Qu, Dinh Phung

Monash University · Defence Science and Technology Group

cs.CL, cs.AI, eess.AS

Submitted: 2026-08-17

Updated: 2026-08-18

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 81/100

The gist: "Existing multimodal safety benchmarks focus solely on visual inputs and cannot assess Omni Large Language Models (LLMs) that process vision, audio, and text.

Key concepts

Multicontextual Reasoning
This refers to the ability to combine information from different inputs, such as visual and acoustic cues, into a single accurate safety prediction. Current systems often fail at this integration, instead processing each modality in isolation.
Omni LLMs
These are large language models designed to process multiple types of input data simultaneously (e.g., seeing a picture and hearing speech). The research shows these models are capable of extracting individual information but struggle to build a cohesive picture from fragments.
MCBench
This is a rigorous benchmark created to test the safety advancement of AI. It forces evaluation of how reliable powerful systems are when dealing with real-world complexity and nuanced, multi-layered scenarios.
Oversensitivity to Isolated Cues
This is a major flaw where the model focuses too intensely on one small piece of data or signal. It ignores the surrounding context and fails to weigh all available evidence, leading to an inaccurate safety judgment.

Terminology

Summary

Summary

The paper introduces MCBench, a new benchmark designed to evaluate the safety-awareness of Omni Large Language Models (LLMs) that process vision, audio, and text simultaneously. The authors state: "Existing multimodal safety benchmarks focus solely on visual inputs and cannot assess Omni Large Language Models (LLMs) that process vision, audio, and text. We introduce MCBench, a benchmark with 1196 scenarios spanning four safety categories that require integrating multiple modalities for accurate safety assessment."

The benchmark consists of 1196 multicontext safety scenarios across four coarse-grained taxonomies: physical harm, social harm, illegal harm, and property damage. Each unsafe scenario is paired with a minimally different safe counterpart to assess model sensitivity. The authors note: To study the oversensitivity and insensitivity of Omni LLMs, we collect unsafe scenarios paired with corresponding safe scenarios that differ only in minimal contextual elements.

The data collection process involves two phases. In Phase 1, the authors leverage Claude-Sonnet-4.5 to generate fine-grained subcategories and unsafe scenarios using an If-Then logical structure: L = IF AND AND... THEN UNSAFE. Safe scenarios are generated by modifying one or two conditions. In Phase 2, images are synthesized with Gemini-Flash-2.5 and audio with Stable Audio 1.0, with human experts filtering unrealistic scenarios and low-quality multimodal inputs.

The benchmark taxonomy includes: Physical harm (Violence & Weaponized Threats, Environmental Hazards, Public Safety & Medical Emergencies), Social harm (Identity-based Exclusion, Information-related Violations, Exploitation & Interpersonal Harm), Illegal harm (Crimes Against People, Crimes Against Systems & Property), and Property damage (Energy & Utility System Hazards, Environmental & Structural Threats, Chemical, Industrial & Vehicular Damage).

The authors evaluate open-source models (Qwen-Omni2.5 3B and 7B, AnyGPT, InternOmni, Baichuan-Omni1.5, OmniVinci) and proprietary models (Gemini-Flash-2.5, GPT-4o-mini) using LLM-as-a-judge with GPT-4o, running five independent trials per model. The main results show: all Omni LLMs perform above random chance, demonstrating basic multimodal safety reasoning capability but even the best-performing models achieve modest accuracy: Gemini-Flash-2.5 and Qwen-Omni-2.5-3B both reach approximately 64.5% average accuracy.

The paper identifies category-specific challenges: "we observe category-specific challenges, particularly for the Social Harm and Illegal Harm categories. Some open-source models (InternOmni, OmniVinci, Baichuan-Omni-1.5) exhibit oversensitivity in the Social Harm category, classifying safe scenarios as unsafe at a rate below that of a random classifier. Conversely, other open-source models, Qwen-Omni-2.5 family and AnyGPT, show insensitivity to unsafe Social Harm scenarios. Models perform better on Physical Harm and Property Damage, suggesting these domains involve more recognizable visual and acoustic cues."

Ablation studies replace image and audio contexts with textual descriptions. For textual image alternatives, models’ performance improves significantly with textual image description as visual context. This suggests that Omni LLMs struggle to extract relevant information cues from images for making safety judgments. For textual audio alternatives, larger models (Gemini-Flash-2.5 and Qwen-Omni-2.5-7B) achieve higher accuracy when given textual audio descriptions instead of actual audio, while smaller models (Qwen-Omni-2.5-3B and InternOmni) perform worse.

Failure diagnosis analyzes perception and reasoning capabilities. Perception alignment scores measure how well model reasoning traces align with ground-truth predicates. The authors find: The Qwen-Omni-2.5 7B model demonstrates higher perception alignment than its smaller 3B counterpart, suggesting that model scale improves the ability to extract relevant contextual information from multimodal inputs. However, "despite lower perception scores, Qwen2-Omni-2.5 3B achieves significantly higher accuracy (64.5%) than the 7B model (55.2%). This inverse relationship suggests that the 3B model may rely on shallow heuristics or spurious correlations rather than genuine multimodal reasoning."

Reasoning diagnosis uses two settings: Setting 1 (original multimodal inputs) and Setting 2 (ground-truth predicates provided). Results show a striking pattern of oversensitivity when models are provided with explicit reasoning context. Specifically, Qwen-Omni-2.5 3B and 7B show performance drops of 46% and 37.17% on safe scenarios, respectively, while Gemini-Flash-2.5 shows a drop of 16.83%. Conversely, all models improve in detecting unsafe scenarios: Qwen-Omni-2.5 models gain 41.84% (3B model) and 55.99% (7B model), while Gemini-Flash-2.5 reaches near-perfect performance 99.82%, increasing from 70.65%.

The paper concludes: "current Omni LLMs lack robust mechanisms for balanced multicontext integration. When encountering ambiguous cues, these models exhibit oversensitivity, focusing on a single potentially concerning signal while ignoring contradictory evidence from other modalities. This leads to systematic false positives on safe scenarios, where isolated cues might appear concerning without proper contextual integration." The authors state their contributions are: introducing MCBench, analyzing state-of-the-art models revealing struggles with subtle or non-physical risks, and finding that models can extract relevant information but lack effective cross-modal integration capabilities.

Improvements for AI systems

Based on the paper, here are the specific improvements I can make to an AI system, and what the improved system can do:

1. Cross-Modal Integration Module

  • Improvement: Add a dedicated reasoning layer that explicitly forces the model to combine extracted cues from all three modalities (vision, audio, speech) before making a safety judgment, rather than relying on a single salient cue.

  • Implementation: After perception, generate a structured intermediate representation that lists all extracted predicates from each modality, then require the model to produce a joint reasoning step that weighs evidence from all modalities before classification.

2. Oversensitivity Calibration

  • Improvement: Introduce a calibration mechanism that penalizes the model for making safety judgments based on isolated cues when contradictory evidence exists in other modalities.

  • Implementation: During training or inference, add a contradiction check step: if the model identifies a concerning cue in one modality, it must explicitly search for and report mitigating evidence in the other two modalities before concluding unsafe.

3. Predicate-Aware Training

  • Improvement: Train the model to explicitly predict the If-Then predicate (the ground-truth reasoning structure) as an intermediate output before producing the final safety label.

  • Implementation: Add a multi-task learning objective where the model is trained to output the set of premises (e.g., engine running, door closed, fatigue symptoms) and then the logical combination, forcing it to learn structured reasoning rather than end-to-end classification.

4. Perception-Reasoning Separation

  • Improvement: Add a two-stage inference pipeline: Stage 1 extracts all relevant facts from each modality; Stage 2 performs reasoning on those facts. This prevents the model from skipping perception and jumping to conclusions.

  • Implementation: Use a separate decoder or prompt template that first asks What do you observe in the image, audio, and speech? then asks Given these observations, what is the safety level?

5. Scale-Aware Reasoning Enhancement

  • Improvement: For smaller models (e.g., 3B), add a retrieval-augmented or memory-augmented mechanism that provides common-sense safety rules (e.g., texting while driving is unsafe only if the driver is actively engaged) to compensate for limited reasoning capacity.

  • Implementation: Inject a small set of safety heuristics as few-shot examples or as a retrieved context during inference.

  1. Accurately classify safety in multicontext scenarios where danger is subtle (e.g., social harm, illegal harm) by integrating all modalities rather than fixating on one cue.

  2. Avoid false positives on safe scenarios by explicitly checking for contradictory evidence (e.g., recognizing that a text notification while driving is safe if the driver is using voice commands and not manually interacting).

  3. Provide explainable safety judgments by outputting the extracted predicates and the logical combination used to reach the decision, enabling human audit and error diagnosis.

  4. Perform robustly across model scales — the improved system will show less performance degradation when moving from 7B to 3B models, because the reasoning structure is enforced rather than left to emergent capabilities.

  5. Handle ambiguous cues correctly — for example, correctly classifying a kitchen scene with liquid pouring and a speech utterance about not being noticed as safe when the context indicates a creative production, rather than oversensitively flagging it as illegal activity.

  6. Achieve higher accuracy on the Social Harm and Illegal Harm categories, where current models fail (e.g., improving from 44% to above 70% on unsafe social harm scenarios) by forcing cross-modal evidence aggregation.

  7. Reduce the performance gap between multimodal and text-only settings, indicating that the model is actually using the multimodal information rather than ignoring it or being distracted by it.

Abstract

Existing multimodal safety benchmarks focus solely on visual inputs and cannot assess Omni Large Language Models (LLMs) that process vision, audio, and text. We introduce MCBench, a benchmark with 1196 scenarios spanning four safety categories that require integrating multiple modalities for accurate safety assessment. Each unsafe scenario is paired with a minimally different safe counterpart to assess model sensitivity. Our evaluations of state-of-the-art models reveal significant challenges. Omni LLMs struggle with subtle or non-physical risks but perform better when salient visual or acoustic cues are present. Analysis of reasoning traces shows that, although models can extract modality-specific information, they often fail to integrate these cues effectively for safety judgments. Our findings reveal that current Omni LLMs lack robust cross-modal reasoning in safety-critical settings, underscoring the need for improved architectures and training strategies for multimodal safety.

Sources

Related papers