The Attention Triangle in Audio-Video Models
cs.AI
Submitted: 2026-09-03
Updated: 2026-09-03
License: http://creativecommons.org/licenses/by/4.0/
The gist: Audio-video diffusion models rely on cross-modal attention to coordinate text, sound, and visual content, yet this same mechanism can introduce subtle and systematic semantic leakage.
Terminology
Abstract
Audio-video diffusion models rely on cross-modal attention to coordinate text, sound, and visual content, yet this same mechanism can introduce subtle and systematic semantic leakage. We study these models by probing and analyzing the ``attention triangle,'' comprising the three cross-attention edges connecting the text, audio, and video streams, and examine how semantic information is routed across modalities during generation. Our analysis reveals that routing along the audio-video edge is bidirectional: audio can influence video generation, while video can influence audio generation. This edge is shaped by biases encoded in the model's parameters and emerges as a major contributor to leakage: when prompts are in tension with learned priors, cross-modal interactions may override the intended conditioning and reroute semantics toward visually canonical but incorrect outcomes. These effects suggest that semantic artifacts arise not merely from attention spreading beyond its intended target, but from structured, bias-driven interactions along specific pathways. Building on this perspective, we extract attention-derived signals that expose how semantics are distributed and grounded across modalities, and use them as a diagnostic tool to both analyze and deliberately incur leakage under controlled conditions. This enables us to probe the internal dynamics of cross-modal routing and isolate the role of individual interactions. We further leverage these signals to guide inference-time interventions that encourage more consistent cross-modal alignment. Extensive experiments support our analysis and demonstrate improved semantic grounding while preserving generation quality.
Sources
- Do Joint Audio-Video Generation Models Understand Physics?
- LTX-2: Efficient Joint Audio-Visual Foundation Model
- Imagen Video: High Definition Video Generation with Diffusion Models
- VA-Judger: Reward Modeling from Human Preference Feedback for Joint Video-Audio Generation
- Image Generation from Contextually-Contradictory Prompts
- SAM 3: Segment Anything with Concepts
- MMControl: Unified Multi-Modal Control for Joint Audio-Video Generation
- SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text
- Improving Joint Audio-Video Generation with Cross-Modal Context Learning
- DeLeaker: Dynamic Inference-Time Reweighting For Semantic Leakage Mitigation in Text-to-Image Models
- Qwen3-Omni Technical Report
- CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection