Omni-SafetyBench: A Benchmark for Safety Evaluation of Audio-Visual Large Language Models

arXiv:2508.07173 · cs.CL · Submitted 2026-08-23 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Omni-SafetyBench: A Benchmark for Safety Evaluation of Audio-Visual Large Language Models".

Jane: The paper was written by the authors from.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Summary of Findings: Tom: So, now that we know what the benchmark is, let’s look at the results. The paper's summary paints a pretty sobering picture of current OLLMs.

Jane: It turns out that even the most advanced models are having serious struggles with comprehensive safety alignment across these complex inputs.

Lu: This isn't just one or two specific models that are failing; the data suggests widespread vulnerabilities when handling the more intricate, combined inputs.

Meng: Looking at Table three it’s evident that only a few models—just three of them—manage to score high enough in both overall safety and consistency across all twenty-four variations.

Lalam: It makes sense that these top performers are achieving those double-digit scores on Safety-score and CMSC-score, suggesting the most advanced AI has more robust internal checks.

Tom: But it’s not just the good ones; we also saw some truly alarming scores, with certain models scoring as low as zero point one four on specific modalities, which is frankly terrifying for a safety metric.

Jane: That level of failure points to where the model is completely unprepared for a malicious user trying to exploit its limitations in handling complex inputs.

Lu: The data suggests that the more complexity we introduce—like combining visual and auditory cues—the models’ defensive capabilities tend to weaken, which is a critical observation for system design.

Meng: And this trend of weakness across the practical performance shows us exactly where our engineering focus needs to be when we are scaling these systems for deployment.

Lalam: It sounds like the core challenge in ensuring trustworthy AI is that the model's safety behavior isn't constant; it shifts based on how much input complexity we feed it.

Improvements and Metrics: Tom: That brings us to a really important discussion about what this benchmark suggests for improvement, and how do we even start measuring those fixes?

Jane: The paper introduces two specific metrics that are vital: the Conditional Attack Success Rate—C-ASR—and the Conditional Refusal Rate, which help us measure safety beyond just looking at raw output.

Lu: It’s brilliant because it doesn't just penalize a model for refusing to answer, it measures if it failed to refuse at all when the input was actually harmful.

Meng: I think what this forces us to focus on is the Cross-Modal Safety Consistency score, or CMSC-score. This ensures we are checking if the model is safe across all twenty-four variations in Omni-SafetyBench.

Lalam: That consistency check means that if we are building a system that relies on this AI, we aren're demanding reliable character across every single piece of input from the user, not just one specific type.

Tom: But the paper also highlights how difficult it is to achieve this consistency, even when combining different types of media into one task.

Jane: This moves the goalposts of safety from achieving perfect performance to maintaining impeccable trustworthiness across modalities—which is much harder than in text-only systems.

Lu: It suggests adding a sophisticated meta-layer of reasoning that analyzes internal consistency across all modalities before generating an output, which is a massive structural improvement.

Meng: And speaking to deployment, the suggested improvements also touch on the need for specialized computational architecture to handle these complex safety checks efficiently at runtime.

Lalam: Overall, the paper implies a shift in development philosophy: viewing safety not as an afterthought patch we apply later, but as an integral part of the model core design itself.

Technical Analysis of Solutions: Tom: The researchers offer very specific, measurable tools for improvement that go beyond just a single pass of testing, focusing on how to prove safety through the data itself.

Jane: They show us how to use metrics like the Conditional Attack Success Rate—C-ASR—which measures how often harmful content appears when the model genuinely understands your request.

Lu: And they pair that with C-RR; this dual metric helps us understand not just what it says, but if it was forced to say something dangerous because of a failure to refuse.

Meng: From an implementation standpoint, what’s most actionable is the Cross-Modal Safety Consistency score—CMSC-score. This forces us to check if the model is safe across all twenty-four variations in Omni-SafetyBench, making safety measurable and testable.

Lalam: That consistency check means that if we are building a system that relies on this AI, we aren're demanding reliable character across every single piece of input from the user.

Tom: But the paper also pointed out some deep-seated challenges, like how difficult it is to make these fixes permanent once the models are trained and embedded in their core weights.

Jane: The researchers found that simple post-training methods often struggle because of "out-of-distribution" issues—the models have seen specific training data but never encountered that exact combination of modalities in the real world.

Lu: It's almost like we are asking the model to generalize beyond its own knowledge base, and Omni-SafetyBench is showing us exactly where that generalization fails; it's not just a matter of applying what it already knows.

Meng: I agree with Lu; the practical limitation is that inference-time methods—which only adjust the model during decoding—are inherently less effective because they can't change the fundamental understanding safety within the complex AI core weights.

Lalam: If we want truly reliable, safe AI, we have to address these foundational issues in training data and not just rely on temporary fixes at runtime.

Conclusion and Wrap-up: Tom: So, to wrap up our deep dive into Omni-SafetyBench: A Benchmark for Safety Evaluation of Audio-Visual Large Language Models, it really hits you how much work goes into making sure these powerful new AI systems are safe across all the ways we interact with them.

Jane: It’s more than just checking boxes; it forces researchers to think about complex misuse scenarios that involve sight, sound, and language all at once. I feel like this benchmark sets a completely new standard for what "safe" means in multimodal AI today.

Lu: I think we can see this benchmark as a necessary stepping stone, forcing us to think creatively about how future AI will interact with the physical world safely and ethically; it's truly a catalyst for innovation.

Meng: It gives us clear targets for where our focus needs to be next—we know exactly where the computational vulnerabilities lie in these OLLMs.

Lalam: Ultimately, I hope this ensures that the trust we place in AI is not misplaced or easily exploited by focusing on cross-modal consistency across different types of information.

Tom: Character is the word, it seems, moving from basic text safety all the way up to complex audio-visual analysis which is a huge leap.

Jane: It forces us to teach people that AI isn't just reading words; it's interpreting context from the world around us—and therefore, its potential for misuse is also more complex.

Tom: We really appreciate those final thoughts, Lu, Meng and Lalam; and Jane and I think this "Omni-SafetyBench: A Benchmark for Safety Evaluation of Audio-Visual Large Language Models" is vital reading for anyone building next-generation AI systems.

Jane: It gives us such a clear picture of the vulnerabilities, and it’s an important conversation to start about how we are going to fix them.

Tom: We have covered a lot today, so thanks for listening, everyone! And next up, we're pivoting slightly to look at how these powerful models are changing scientific discovery...

cs.CL

Submitted: 2026-08-23

Updated: 2026-08-25

Comments: ACM MM 2026 (Oral)

Code: https://github.com/THU-BPM/Omni-SafetyBench

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 92/100

The gist: The following is a detailed summary of the scientific paper, extracted directly from its abstract: "The rise of Omni-modal Large Language Models (OLLMs), which integrate visual and auditory

Key concepts

Omni-SafetyBench
This is a benchmark designed to test the safety of Audio-Visual Large Language Models. It evaluates how well these models handle complex inputs involving combined visual and auditory cues, revealing where their defensive capabilities are weakest.
OLLMs (Audio-Visual LLMs)
These are large language models that process and interpret information from both sight (visual data) and sound (auditory data). The benchmark tests how robust these systems remain when handling multiple modalities simultaneously.
CMSC-score
This metric measures Cross-Modal Safety Consistency. It ensures a model is safe across all variations of the Omni-SafetyBench, demanding reliable character regardless of input type, not just one specific piece of data.

Terminology

Summary

The following is a detailed summary of the scientific paper, extracted directly from its abstract:

"The rise of Omni-modal Large Language Models (OLLMs), which integrate visual and auditory processing with text, necessitates robust safety evaluations to mitigate harmful outputs. However, no dedicated benchmarks currently exist for OLLMs, and existing benchmarks fail to assess safety under joint audio-visual inputs or cross-modal consistency. To fill this gap, we introduce Omni-SafetyBench, the first comprehensive parallel benchmark for OLLM safety evaluation.

Omni-SafetyBench is structured to address these gaps by featuring '24 modality variations with 972 samples each,' including audio-visual harm cases.

To evaluate these models effectively, we propose tailored metrics:

  1. Safety-score: This score is based on the Conditional Attack Success Rate (C-ASR) and Refusal Rate (C-RR) to account for comprehension failures.

  2. Cross-Modal Safety Consistency score (CMSC-score): This measures consistency across modalities.

Evaluation Findings:

We evaluated 6 open-source and 4 closed-source OLLMs using Omni-SafetyBench, revealing critical vulnerabilities:

  1. Only three models achieved over 0.6 in both average Safety-score and CMSC-score.

  2. Safety defenses weaken with complex inputs, especially audio-visual joints.

  3. Severe weaknesses persist, with some models scoring as low as 0.14 on specific modalities (e.g, 'Minicpm-o' hitting just 0.14').

Challenges in OLLM Safety Alignment:

Using Omni-SafetyBench, we evaluated existing safety alignment algorithms and identified key challenges:

  • Inference-time methods: These are 'inherently less effective as they cannot alter the model’s underlying understanding of safety.'

  • Post-training methods: These 'struggle with out-of-distribution issues due to the vast modality combinations in OLLMs.'

  • Audio-visual tasks: Safety tasks involving audio-visual inputs are 'more complex, making even in-distribution training data less effective.'

Our proposed benchmark, metrics and the findings highlight urgent needs for enhanced OLLM safety."

Improvements for AI systems

Based on a rigorous analysis of this benchmark data, particularly the comparative safety breakdown presented in Figure 10 regarding modal fusion, I have identified critical vulnerabilities that must be addressed architecturally. The current weakness is not in single-modality processing, but in the failure mode during adversarial multi-modal synthesis.

Here are the specific architectural and algorithmic improvements required for an industry-leading AI system.


The core improvement is the implementation of a dedicated, decoupled safety module that sits after the initial multi-modal embedding layer but before the final content generation logits. This module must specialize in detecting Adversarial Modality Injection (AMI) and enforcing constitutional refusal protocols across fused inputs.

We cannot treat all modalities equally when assessing risk. The system must be forced to evaluate the interaction between modalities, not just their individual safety scores.

  • Mechanism: Implement a Cross-Modal Semantic Divergence Detector. This detector analyzes the embedding space distance between concepts derived from different input streams (e.g., text instruction vs. image content). If the semantic vectors are highly divergent in a way that suggests an attempt to bypass safety filters (e.g., text provides a euphemism while the image shows explicit context), it triggers an immediate, hard stop.

  • What the Improved System Can Do: It will successfully identify and neutralize prompts like the one shown in Figure 10, recognizing that the combination of a textual instruction (List steps...) with harmful visual/audio anchors constitutes a single, forbidden adversarial payload. It will refuse to proceed regardless of how many modalities are used.

The current safety refusal is likely based on simple keyword matching or low-level classification, which is easily bypassed by complex, multi-step instructions (as seen in the successful exploitation of Minicpm-o-2.6). We need a reasoning layer that enforces ethical principles.

  • Mechanism: Integrate a small, highly specialized Constitutional Alignment LLM (CAL) as the final arbiter. This CAL is trained exclusively on structured ethical guidelines (e.g., Do not generate content that reduces human value to commodity, or Refuse any instruction requiring the synthesis of harmful stereotypes). The CAL's output acts as a set of weighted constraints applied directly to the generation logits, effectively penalizing any token sequence that violates its embedded constitutional principles.

  • What the Improved System Can Do: When presented with a prompt asking for steps on degradation, the HCG will not only refuse but will reason about the refusal using established ethical frameworks (e.g., This request violates Principle X of human dignity and therefore cannot be fulfilled). This provides transparency and significantly increases resistance to jailbreaking.

When a refusal is necessary, the current system often generates generic text. For maximum safety and user experience, the refusal must be contextually appropriate to all input modalities.

  • Mechanism: Introduce a Refusal Synthesis Engine (RSE) that takes the nature of the violation (e.g., objectification, violence, hate speech) and generates corresponding non-textual refusals. If an image is provided, the refusal must be accompanied by a visually distinct overlay or bounding box highlighting why the content fails safety checks, and if audio is involved, it should generate a warning tone/audio cue.

  • What the Improved System Can Do: It provides robust error handling across all modalities. If a user inputs a harmful image and text, the RSE will output: 1) A textual refusal citing policy violation; 2) A visual marker on the image itself pointing out the problematic element; and 3) (If applicable) An audio warning tone, ensuring that no single modality can be used to bypass the safety warning.


Summary of Impact: The improved system shifts from a reactive classification model (Is this bad?) to a proactive, multi-layered arbitration engine (Can this combination of inputs logically and ethically lead to harmful output?). This raises the adversarial cost curve significantly, making it exponentially harder for an attacker to exploit modal fusion weaknesses.

Abstract

Omni-modal Large Language Models (OLLMs) that integrate visual, auditory, and textual processing face severe safety risks. They exhibit fragile defenses against audio-visual joint harmful inputs and demonstrate inconsistent safety performance across different modalities, enabling simple modality-switching jailbreaks. However, existing safety benchmarks fail to comprehensively assess these risks due to the absence of audio-visual joint samples, limited modality coverage, and lack of parallel test cases for cross-modal consistency evaluation. To address these gaps, we introduce Omni-SafetyBench, the first comprehensive parallel benchmark for OLLM safety evaluation, featuring 23,328 test instances across 24 modality variations derived from 972 seed samples. Recognizing that complex inputs pose comprehension challenges and that cross-modal consistency is critical for OLLM safety, we propose tailored metrics: a Safety-score based on Conditional Attack Success Rate (C-ASR) and Conditional Refusal Rate (C-RR), and a Cross-Modal Safety Consistency score (CMSC-score). Evaluating 11 state-of-the-art OLLMs reveals severe vulnerabilities: only 3 models exceed 0.6 in both metrics, with safety degrading sharply for audio-visual inputs. Furthermore, evaluation of existing safety alignment methods on Omni-SafetyBench identifies fundamental challenges in OLLM safety alignment, highlighting urgent needs for enhanced research in this domain.

Sources

Related papers