Multi2AV-Safety: Benchmarking Safety in Multimodal-to-Audio-Video Generation
cs.AI
Submitted: 2026-08-27
Updated: 2026-08-27
License: http://creativecommons.org/licenses/by/4.0/
The gist: Audio-video generation is rapidly moving from prompt-driven synthesis toward multimodal conditioning, where text, images, audio, and video can jointly shape the generated output.
Terminology
Abstract
Audio-video generation is rapidly moving from prompt-driven synthesis toward multimodal conditioning, where text, images, audio, and video can jointly shape the generated output. This shift changes the nature of safety evaluation: harmful intent may no longer reside in any single input, but instead emerge from how otherwise benign or weakly harmful conditions interact across modalities and time. Existing safety benchmarks, however, remain largely prompt-centric or tied to fixed conditioning interfaces, leaving such compositional risks difficult to study systematically. To bridge this gap, we introduce Multi2AV-Safety, the first safety benchmark, to the best of our knowledge, to cover all 11 non-singleton T/I/A/V conditioning configurations for audio-video generation, comprising 11,024 attack instances. Evaluation on Multi2AV-Safety reveals systematic weaknesses in representative multimodal safety guards across attack mechanisms and harm-evidence structures. Our evaluation reveals two complementary failure modes: harmful semantics can emerge from the combination of individually benign inputs, while explicit harmful cues can become harder to detect when mixed with benign multimodal context. Together, these results identify compositional risk perception as a central capability gap in safeguarding multimodal-conditioned audio-video generation: current safety guards fail to reliably integrate safety evidence across modalities and time, even when all conditioning inputs are observable. The dataset will be publicly released in October 2026.
Sources
- LTX-2: Efficient Joint Audio-Visual Foundation Model
- Speed by Simplicity: A Single-Stream Architecture for Fast Audio-Video Generative Foundation Model
- ConceptGuard: Proactive Safety in Text-and-Image-to-Video Generation through Multimodal Risk Detection
- SafeGen-Bench: Benchmarking Safety in Image-Conditioned Text-to-Video Generation
- Omni-SafetyBench: A Benchmark for Safety Evaluation of Audio-Visual Large Language Models
- UniSAFE: A Comprehensive Benchmark for Safety Evaluation of Unified Multimodal Models
- HiDream-O1-Image: A Natively Unified Image Generative Foundation Model with Pixel-level Unified Transformer
- CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models
- Harnessing LLM to Attack LLM-Guarded Text-to-Image Models
- Reason2Attack: Jailbreaking Text-to-Image Models via LLM Reasoning
- Contextual Image Attack: How Visual Context Exposes Multimodal Safety Vulnerabilities
- VPA-Guard: Defending and Benchmarking Image-to-Video Generation Against Visual Prompt Attacks
- GenBreak: Red Teaming Text-to-Image Generators Using Large Language Models
- Llama Guard 3 Vision: Safeguarding Human-AI Image Understanding Conversations
- Qwen3-Omni Technical Report
- GuardReasoner-Omni: A Reasoning-based Multi-modal Guardrail for Text, Image, Video, and Audio
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection