mAVE: A Watermark for Joint Audio-Visual Generation Models

arXiv:2603.07090 · cs.CR, cs.AI, cs.CV · Submitted 2026-03-07 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.

Nadia: Today's paper: "mAVE: A Watermark for Joint Audio-Visual Generation Models".

Elias: The gist The proposed mAVE framework is the first watermarking strategy natively designed for joint architectures,

Nadia: First, who's behind it and why it matters.

Paper summary: Nadia: So we're looking at this paper called mAVE: A Watermark for Joint Audio-Visual Generation Models, and the main idea is that existing protection methods don't work well when you have both audio and video going at once.

Elias: Exactly. The authors point out that current techniques treat audio and video like separate things, which creates a problem they call the Binding Vulnerability. This vulnerability lets an adversary swap out authentic audio for something malicious deepfake while keeping the video's watermark intact because detectors check them separately.

Priya: That sounds like a real risk because if you have two independent checks, swapping one element can fool both of them simultaneously, leading to false authentication of harmful content >

Nadia: Right. And mAVE tackles this by designing a strategy that is built right into the initialization process so the audio and video are cryptographically bound together from the start. They claim this creates a formal Legitimate Entanglement Manifold >

Elias: They achieve this by securely entangling the audio latent to the video latent using a specific function, z a equals f(z v), which forces any swap to break that functional dependency >

Priya: So what they're saying is that instead of just slapping a watermark on each modality separately, they are building a shared space where the audio and video are mathematically linked in a way that makes it much harder to tamper with both at once >

Nadia: It’s about defining this entanglement through something called Inverse Transform Sampling, which restricts the whole joint generation process to stay on this entangled manifold rather than just reducing the dimensionality of each part separately >

Elias: They set up a discrete entangled geometry where the audio grid is bound to the video grid by embedding a hash digest of the video bits into that audio grid at specific indices >

Priya: But what does that actually look like in practice for someone using this kind of model? Is it just more math, or does it actually change how you generate content >

Nadia: The paper claims they do this without needing any extra fine-tuning or adding artifacts to the final output, which is a big deal because most watermarking methods require some sort of adjustment after the model is trained >

Elias: They give two main guarantees about this framework: first, Performance-Losslessness. They prove that under their specific tests, the watermarked latent z s follows the same distribution as standard Gaussian initialization >

Priya: That means if you run a clean model and then apply mAVE's method, the resulting quality of the video or audio shouldn't degrade noticeably when you look at it >

Nadia: And they also have this Security Bound. This is where they show that the probability of an adversary successfully swapping something passes their check drops off exponentially with N, which is a measure of model size >

Elias: They give a specific number for this: for a default configuration with N equals one hundred twenty-eight and taubind set to zero point eight, the evasion probability is shown to be less than nine point eight six times ten to the minus eleven >

Priya: That’s a very small probability, which suggests that this method offers a strong defense against someone trying to use deepfakes maliciously >

Nadia: So for someone just listening, what does this mean for their day-to-day experience when they are using these kinds of generative tools >

Elias: It means that instead of relying on separate watermarks that can be easily bypassed, the system itself enforces a cryptographic link between the audio and video components during creation >

Priya: It suggests that the real value here isn't just in having a watermark, but in how you structure the entire generation process to make tampering with one part inherently detectable across both modalities >

Nadia: We’re going to take a quick break and then we’ll talk about what this whole mAVE thing means for the future of protecting digital media >

Conclusion: Nadia: So we're wrapping up on mAVE, which is this new watermarking strategy designed specifically for models that handle both audio and video together.

Elias: Yeah, the core idea is that it’s the first one built from scratch to cryptographically link the audio and video latents right at the very start of the generation process.

Priya: What that actually means for us is moving away from treating audio and video like separate files you can just slap a label on.

Nadia: Right, because existing methods fail when you have those two modalities interacting, creating this binding vulnerability where an adversary can swap audio for a deepfake while keeping the video watermark intact.

Elias: Exactly. They're exploiting that mismatch by using independent detectors, which just don't catch the manipulation because they aren't looking at the combined system.

Priya: And mAVE fixes that by creating this legitimate entanglement manifold, essentially forcing the audio and video to stay coupled in a way that’s mathematically enforced.

Nadia: It uses a two-step process involving hashing bits from the video into the audio grid, which then gets diffused back out using inverse transform sampling.

Elias: That’s how they construct this discrete geometry where the hash digest of the video is embedded directly into specific spots in that audio structure.

Priya: And they use a session key derived from a secret payload and a prompt to randomize those bits, mapping them onto the continuous latent space with an inverse probability integral transform.

Nadia: The theoretical guarantees are pretty strong here too, showing performance-losslessness under certain tests and an exponential security bound against swap attacks.

Elias: That exponential bound is what’s interesting for me—it suggests that the chance of someone successfully swapping a pair drops off really fast as the model gets bigger.

Priya: From a measurement side, they show that in experiments on models like LTXtwo and MOVA, mAVE actually outperforms just using separate watermarks.

Nadia: So to sum up, mAVE is this training-free method that embeds cryptographic binding directly into the initialization to guarantee performance and security.

Elias: It’s a lot of math behind it, but they’re proving that you can enforce this link without sacrificing generation quality or introducing noticeable artifacts.

Priya: The limitation they mentioned is that there's some deterministic drift caused by how the ODE discretization happens at a small time step.

Nadia: That means while the theory is solid, we still have to be aware of that tiny bit of numerical error when using it in practice.

Elias: So, mAVE shifts the focus from post-generation detection to building security into the very architecture of how these joint models start up.

Priya: It opens up a new way for researchers to think about protecting multi-modal generative content by focusing on structural integrity rather than just adding stickers onto it.

School of Software, Tsinghua University

cs.CR, cs.AI, cs.CV

Submitted: 2026-03-07

Updated: 2026-10-08

License: http://creativecommons.org/licenses/by-nc-sa/4.0/

Importance score: 91/100

The gist: The gist The proposed mAVE framework is the first watermarking strategy natively designed for joint architectures, cryptographically binding audio and video latents at initialization to guarantee

Key concepts

Binding Vulnerability
This is a security gap in current protection methods where audio and video are treated separately, allowing attackers to swap authentic audio with malicious deepfakes without detection. Current detectors fail because they check modalities independently, leading to false authentication of manipulated content.
Legitimate Entanglement Manifold
mAVE constructs a specific mathematical space where the entangled audio and video latents must reside. This manifold is defined by sparse cryptographic constraints, forcing the joint generation process to adhere to this structure, which ensures that only cryptographically linked pairs can be generated.
Cryptographic Binding at Initialization
This technique securely entangles the audio latent with the video latent right when generation starts. It involves embedding a hash digest of the video bits into specific indices of the audio grid, creating an unbreakable link between the two modalities from their very beginning.

Terminology

Summary

The gist The proposed mAVE framework is the first watermarking strategy natively designed for joint architectures, cryptographically binding audio and video latents at initialization to guarantee performance-losslessness and provide an exponential security bound against Swap Attacks

The Problem

Existing protection mechanisms suffer from a fundamental architectural mismatch by treating modalities as decoupled entities, exposing a critical Binding Vulnerability Adversaries exploit this via Swap Attacks by replacing authentic audio with malicious deepfakes while retaining the watermarked video Because current detectors rely on independent verification (V ideowm ∨ Audiowm), they incorrectly authenticate the manipulated content, falsely attributing harmful media to the original vendor and severely damaging their reputation Crucially, elevating the detection criteria to a logical conjunction (V ideowm ∧ Audiowm) fails to mitigate this threat

How it works

The mAVE framework enforces a Cryptographic Binding at initialization by securely entangling the audio latent to the video latent (za = f(zv)), constructing a Legitimate Entanglement Manifold via Inverse Transform Sampling This method restricts the joint generation process to a cryptographically entangled manifold, where the Authentic Manifold M is defined by sparse cryptographic constraints rather than dimensionality reduction

The process involves a two-step pipeline for initial noise:

  1. Constructing the discrete entangled geometry (Sec. 4.1) where the Audio Grid (Ba) is cryptographically bound to the Video Grid (Bv) via a hash digest This coupling is achieved by embedding a digest of the video bits, hv = SHA-256(Bv), into the audio grid at indices Ibind

  2. Embedding via Inverse Transform Sampling (Sec. 4.2) where the discrete watermark bits are replicated block-wise to form a diffused tensor Bdif f matching the target latent dimensions This is followed by Watermark Randomization using a session key Ksess derived from the secret payload m and prompt P Finally, the randomized binary map Mrand is mapped to the continuous Gaussian latent z using the Inverse Probability Integral Transform (Eq. 6)

Theoretical Guarantees

The framework provides rigorous theoretical guarantees for its performance and security Theorem 1 proves Performance-Losslessness under chosen watermark tests by showing that the watermarked latent z s follows the same distribution as the standard Gaussian initialization The proof demonstrates that for any polynomial-time tester A and key Ksess ← KeyGen(1ρ), the performance loss is negligible Theorem 2 provides an exponential security bound against Swap Attacks, showing that the probability of a swapped pair passing the binding check decays exponentially with N Specifically, for a default configuration N = 128 and τbind = 0.8, the evasion probability is upper-bounded by Pfp < 9.86 × 10−11

Performance and Detection

Experiments on state-of-the-art models (LTX2 [9] and MOVA [29]) demonstrate that mAVE significantly outperforms naive combinations of unimodal watermarks The framework achieves superior detection accuracy and robust separation bounds against manipulation while maintaining original generation quality In terms of fidelity, mAVE incurs negligible degradation compared to the Clean baseline, with scores statistically indistinguishable from the Clean baseline The extraction cost is identical to VideoShield’s single ODE inversion because it leverages the unified architecture of native bimodal models

Security Analysis

The security analysis confirms that mAVE achieves 99.8% Accuracy against Swap Attacks, which is a substantial improvement over the Weak Baseline's 50% Accuracy The proof for Theorem 2 shows that the binding score Sbind decays exponentially with N This security holds even against adaptive (white-box) adversaries because the objective function is encrypted and any gradient-based attack degenerates to a blind brute-force search over 2N bits The temporal robustness analysis shows that FrameAverage preserves alignment and thus maintains high accuracy, while FrameSwap causes moderate local misalignment with an expected accuracy of approximately (1 − p) · BAclean + p · 0.5

Conclusion

In this work, we identified a critical security gap in the emerging class of joint audio-visual generation models: the Binding Vulnerability, which allows adversaries to decouple and swap modalities without detection To address this, we proposed mAVE (Manifold Audio-Visual Entanglement), the first training-free framework that embeds cryptographic binding directly into the generative initialization By formalizing the watermarking process as a projection onto a Legitimate Entanglement Manifold, mAVE achieves theoretically guaranteed performancelosslessness (Theorem 1) and provides an exponential security bound against Swap Attacks (Theorem 2) Extensive experiments on LTX-2 [9] and MOVA [29] demonstrate that mAVE significantly outperforms unimodal baselines in distinguishing authentic pairs from manipulated ones, while maintaining SOTA generation quality Limitations include deterministic drift due to ODE discretization at t = δt, which prevents raw Bit Accuracy from reaching a theoretical 1.

Improvements for AI systems

  1. textbf Cryptographic Binding in Joint Latent Space Embedding (mAVE): The improved system will embed watermarks by cryptographically binding audio and video latents at initialization without fine-tuning, constructing a Legitimate Entanglement Manifold via Inverse Transform Sampling. This ensures the watermark is intrinsic to the unified generation process, preventing decoupling that leads to a critical Binding Vulnerability.

  2. textbf Robust Swap Attack Defense Against Cross-Session Splicing: The system will achieve an exponential security bound against Swap Attacks, demonstrated by bounding the False Positive Rate (Pfp) at 99.8% Accuracy for mAVE compared to 50% for Weak Baseline (Uncoupled). This allows the AI to confidently reject manipulated content where adversaries retain a vendor’s watermarked video while replacing the authentic audio with a harmful deepfake voiceover.

  3. textbf Efficient, Single-Pass Extraction: The improved system will reduce detection cost by utilizing a single Joint Inversion pass through the bimodal model, which is identical in cost to VideoShield’s video-only inversion while simultaneously recovering both modalities. This eliminates the need for an additional audio encoder, effectively halving the extraction workload compared to baseline combinations.

  4. textbf Performance-Lossless Fidelity: The system guarantees that watermarking does not degrade generation quality, as mAVE is proven performance-lossless under chosen watermark tests, maintaining metrics like CLAP Score and Subject Consistency nearly identical to the Clean baseline. This ensures the AI remains capable of generating high-quality synchronized multimedia while protected.

  5. textbf Adaptive Adversary Resistance: The system offers resilience against white-box adversaries by ensuring that the objective function is encrypted via the server-side secret, meaning an adversary cannot evaluate gradients to bypass the binding check. This forces any attack to degenerate to a blind brute-force search over the 2N possible bit configurations at the binding positions.

Abstract

Watermarking joint audio-visual generation supports vendor copyright protection and content provenance. However, independently valid audio and video watermarks do not establish a shared generation session. An adversary can splice watermarked modalities from different sessions, causing the pair to be mistaken for the vendor's original joint output. We introduce mAVE (Manifold Audio-Visual Entanglement), a training-free watermarking framework that strengthens vendor attribution through session binding in native joint audio-visual diffusion transformers. mAVE separates public record retrieval from secret session authentication: a fixed public index locates the server record, while a randomized payload binds audio bits to a session-keyed video grid through a cryptographic digest. One prompt-conditioned joint inversion supports provider-assisted verification of both modalities against a session record, without modifying generator weights or training auxiliary watermark networks. Our analysis establishes implementation-matched distribution preservation and a full-initialization routing/clipping budget, alongside adaptive session-pool security and stable local-perturbation bounds. Experiments on LTX-2 and MOVA show comparable generation quality. mAVE achieves 99.8% true-positive rate and 0% observed false-positive rate in the evaluated swap test, and retains 99.2% true-positive rate under FrameAvg temporal averaging. Same-prompt and similarity-selected swaps further test session authentication beyond perceptual compatibility.

Sources

Related papers