MARC: Multi-Bit Watermarking for Autoregressive Audio Generation against Codec Attacks
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.
Nadia: Today's paper: "MARC: Multi-Bit Watermarking for Autoregressive Audio Generation against Codec Attacks".
Elias: The gist The proposed MARC framework is a multi-bit generative watermarking method for autoregressive audio generation that integrates intrinsic token representations with confusion patterns obtained through retokenization and multiple codecs…
Nadia: First, who's behind it and why it matters.
Paper summary: Nadia: The main idea here is to create a multi-bit watermarking method specifically for AI-generated audio that can survive those codec attacks, which are really common now. They claim they built something new to solve the issue of zero-bit detection in existing techniques, which means you can find out if a watermark is there but you can't tell what the actual message was.
Elias: Yeah, so it’s about moving from just checking for presence to actually recovering multiple bits of information embedded in the audio, which is pretty difficult because retokenization and codec processing directly corrupt those embedded bits. They propose integrating intrinsic token representations with confusion patterns from both retokenization and multiple codecs to form what they call a codec-aware token-cluster space.
Priya: From a measurement side, I’m interested in how they handle that data corruption; if you're just looking at the raw audio or even the reconstructed version, how do you actually measure that confusion pattern reliably across different compression methods?
Nadia: They address this by building this codec-aware token-cluster space, which is essentially a way to group tokens based on both how they look intrinsically and how they get confused by different codecs. They say this approach helps them build a more robust way to find the correct token cluster, which is key because it moves beyond just looking at one source of information.
Elias: And to handle the zero-bit limitation, MARC uses something called payload-driven cluster scheduling. This lets them embed a multi-bit watermark and then perform detection and decoding on those retokenized observations to recover both if the watermark is there and what the embedded bits actually are.
Priya: That sounds promising for actual recovery, but how does this complex clustering work in practice? How do you define these geometric centers for token relationships when you don't know which transformation will happen later on?
Paper summary: Nadia: They project those generator’s audio-token embeddings into a reduced representation space and fit a Gaussian mixture model to get geometric centers for token relationships, and they say this works independently of which transformations appear in the calibration data. They also accumulate confusion counts by pairing tokens at matching indices over the common sequence length to capture empirical changes across reconstruction and codec channels.
Elias: The codec union is built by combining those matrix of confusion counts from both reconstruction and multiple codecs, where they use an element-wise maximum to keep strong channel-specific confusions, and beta adds weight to the ordinary retokenization. They then get a final codec-aware cluster map, C: Va → K, by combining embedding similarity and normalized codec confusion using that formula A = αS + (one − α)N (MU), where µk is calculated as a weighted combination of the intrinsic and substitution confusion.
Priya: So it’s using both the token structure geometry and the actual observed corruption patterns from compression to define where the signal belongs, which sounds like a lot of data to process. What does this mean for someone just listening to podcasts or watching videos?
Nadia: It means that if you are generating audio with AI, this method aims to put a multi-bit watermark in there and make sure that even after some kind of compression or retokenization happens, you still have a good chance of pulling out the original message. The summary says they achieve an average of ninety-seven point three percent bit extraction accuracy on unmodified watermarked audio and it’s robust against diverse codec attacks <ref:2610.11488#pg1>.
Elias: And for the people who are looking at the technical details, they also introduced a public check value, a pub-key, and a secret auxiliary suffix, the pri-key. Those are used for message consistency checking throughout the verification and recovery protocol without needing any actual cryptographic security assumptions.
Priya: That’s interesting because it means you don't need some complicated key management system just to verify if the watermark is intact, but what about when an attacker tries to tamper with the audio to steal or change those bits?
Paper summary: Nadia: The method also addresses that by using payload-driven cluster scheduling, so they can embed a multi-bit watermark and then perform detection and decoding on retokenized observations to recover both watermark presence and the embedded bits. They even include an extraction and detection step where they evaluate waveform variants and offsets of the watermark-step indices to compute hδ = X t one
ect = kt+δ(m⋆, ret): .
Elias: For the actual decoding part, they do something intensive: they enumerate all sixteen values at every symbol position to compute a hard symbol score Lhardq,a(s) = X t∈Oa jt=q log ηa1
ect = kt(s, ret): + one − ηa. The final payload is initialized by a bit-wise majority voting of the n selected candidates, followed by a local refinement guided by pooled symbol scores and the pub-key.
Priya: So it’s this detailed scoring and voting process that allows them to recover those multiple bits from what looks like corrupted data, rather than just getting a yes or no on whether a watermark exists?
Nadia: Exactly. And for robustness, they showed it can even handle the scenario where an additional watermark is embedded into the released audio without losing the detection and recovery of the original payload. The paper also shows that their full codec-union construction C0 achieves a bit accuracy of ninety-eight point one seven percent and a low FAD of zero point two six two two on clean audio, which is pretty solid compared to other variants they tested.
Elias: And they showed that the full decoder D0 improves both bit accuracy and TPR at FPR=one percent over other decoders under every codec attack. That means their decoding process is more reliable than what was previously tested in the evaluation.
Priya: It seems like a solid piece of research because it tackles the real-world problem of verifying content integrity in an era where audio is everywhere, and they’ve put together a system that actually attempts to recover data rather than just detecting its presence.
Nadia: So, MARC isn't just about finding a watermark; it’s about creating a framework that understands how different AI generation processes and codec transformations affect the token structure so it can recover multi-bit information reliably. That’s what this paper is really focused on in MARC: Multi-Bit Watermarking for Autoregressive Audio Generation against Codec Attacks.
Conclusion: Nadia: So we’re wrapping up on MARC, which is about this new way to watermark AI audio so it survives those codec attacks we talked about earlier, and who wrote this stuff is Nadia, Elias, and Priya.
Elias: Yeah, what you really want to know is what the title itself means because "Multi-Bit Watermarking" sounds like a lot of jargon.
Priya: I think the core idea is that they’re not just putting one tiny bit in there anymore; they’re trying to grab a whole chunk of information so even after compression, you can still pull out what was originally there.
Nadia: Exactly, and what this means for us listeners is that if you're using AI tools to make audio—maybe for voice cloning or music generation—this method aims to make sure the original message isn't completely scrambled beyond recognition.
Elias: From a cryptography standpoint, the authors are building a system that relies on token representations and confusion patterns from multiple codecs to create this specialized space for finding the hidden bits.
Priya: The results show they’re getting around ninety-seven percent accuracy on unmodified audio, which is solid when you have to fight against actual compression artifacts.
Nadia: And the caveat we need to watch is that they showed it can handle additional watermarks being added later without losing the original recovery capability, which suggests a pretty resilient design.
Elias: That resilience relies on how they build this codec-aware clustering, using that mix of intrinsic token data and observed confusion counts from different processing steps.
Priya: It’s really about moving past just detecting if something is there to actually recovering the embedded payload when the audio has been retokenized or compressed multiple times.
Nadia: So, MARC isn't just a theoretical idea; it’s a framework designed to recover multiple bits reliably from AI-generated audio that’s been run through various codecs.
Elias: That kind of multi-bit recovery is what pushes the limits of how much you can hide data in the signal before it becomes completely useless.
Liaoran Xu, Weizhi Liu, Zhaoxia Yin
East China Normal University
cs.CR
Submitted: 2026-10-08
Updated: 2026-10-08
Project page: https://google-research.github.io/seanet/musiclm/examples
License: http://creativecommons.org/licenses/by/4.0/
The gist: The gist The proposed MARC framework is a multi-bit generative watermarking method for autoregressive audio generation that integrates intrinsic token representations with confusion patterns obtained
Key concepts
- Codec-Aware Clustering
- This process projects audio tokens into a smaller space using a Gaussian mixture model to find geometric centers for token relationships. It is designed so that these cluster centers are independent of the specific transformations applied by different codecs, allowing the system to capture channel-specific confusion patterns effectively.
- Codec Union
- The Codec Union is constructed by combining confusion counts derived from both the original reconstruction and multiple codec outputs. By taking the element-wise maximum of these counts and adding a weight for retokenization, it creates a comprehensive map that retains strong channel-specific confusion while also accounting for common retokenization effects.
- Payload Embedding
- The watermark payload is partitioned into three symbols (s0, s1, s2) to be embedded. The target cluster for the watermark is then determined using a formula that incorporates the symbol and various contextual factors like retokenization data and sequence length. This guides the generator to sample tokens from a specific group.
- Extraction and Detection
- Detection involves comparing observed waveform variants against expected watermark-step indices to calculate a score. Decoding uses this score across all 16 possible symbol values at each position to compute a hard symbol score, which is then used for majority voting and local refinement to reconstruct the original multi-bit payload.
Terminology
Summary
The gist The proposed MARC framework is a multi-bit generative watermarking method for autoregressive audio generation that integrates intrinsic token representations with confusion patterns obtained through retokenization and multiple codecs to form a codec-aware token-cluster space, achieving an average of 97.3% bit extraction accuracy on unmodified watermarked audio and robustness against diverse codec attacksMARC: MULTI-BIT WATERMARKING FOR AUTOREGRESSIVE AUDIO GENERATION AGAINST CODEC ATTACKS
How it works
MARC integrates intrinsic token representations with confusion patterns obtained through retokenization and multiple codecs to build a codec-aware token-cluster spaceMARC: MULTI-BIT WATERMARKING FOR AUTOREGRESSIVE AUDIO GENERATION AGAINST CODEC ATTACKS
The framework addresses single-source token grouping by combining intrinsic token representations with confusion patterns induced by retokenization and diverse codecs to build a codec-aware token-cluster spaceMARC: MULTI-BIT WATERMARKING FOR AUTOREGRESSIVE AUDIO GENERATION AGAINST CODEC ATTACKS
To address the zero-bit limitation, MARC uses payload-driven cluster scheduling to embed a multi-bit watermark and performs detection and decoding on retokenized observations to recover both watermark presence and the embedded bitsMARC: MULTI-BIT WATERMARKING FOR AUTOREGRESSIVE AUDIO GENERATION AGAINST CODEC ATTACKS
Codec-Aware Clustering
The method projects the generator’s audio-token embeddings into a reduced representation space and fits a Gaussian mixture model to obtain geometric centers for token relationships independently of which transformations appear in the calibration dataMARC: MULTI-BIT WATERMARKING FOR AUTOREGRESSIVE AUDIO GENERATION AGAINST CODEC ATTACKS
To capture empirical changes across the reconstruction and codec channels, confusion counts are accumulated by pairing tokens at matching indices over the common sequence lengthMARC: MULTI-BIT WATERMARKING FOR AUTOREGRESSIVE AUDIO GENERATION AGAINST CODEC ATTACKS
The codec union is constructed by combining the matrix of confusion counts from the reconstruction and multiple codecs, where the element-wise maximum retains strong channel-specific confusions while beta adds weight to ordinary retokenizationMARC: MULTI-BIT WATERMARKING FOR AUTOREGRESSIVE AUDIO GENERATION AGAINST CODEC ATTACKS
The final codec-aware cluster map C: Va → K is yielded by combining embedding similarity and normalized codec confusion using the formula A = αS + (1 − α)N (MU), where µk = αµGk + (1 − α)µSkMARC: MULTI-BIT WATERMARKING FOR AUTOREGRESSIVE AUDIO GENERATION AGAINST CODEC ATTACKS
Watermark Embedding
The embedded payload m ∈ 0, 12 is partitioned into three symbols s0, s1, s2 ∈ 0,..., 15MARC: MULTI-BIT WATERMARKING FOR AUTOREGRESSIVE AUDIO GENERATION AGAINST CODEC ATTACKS
The target cluster is determined by the formula kt = b(sjt) + H(sjt, rt, lt, jt) mod KMARC: MULTI-BIT WATERMARKING FOR AUTOREGRESSIVE AUDIO GENERATION AGAINST CODEC ATTACKS
The generator samples from qt, which is derived by applying a reweighting rule that assigns non-audio tokens to the remaining groups in proportion to Pt,kMARC: MULTI-BIT WATERMARKING FOR AUTOREGRESSIVE AUDIO GENERATION AGAINST CODEC ATTACKS
Extraction and Detection
The verifier evaluates waveform variants, observation windows, and offsets of the watermark-step indices to compute hδ = X t 1[ect = kt+δ(m⋆, ret)]MARC: MULTI-BIT WATERMARKING FOR AUTOREGRESSIVE AUDIO GENERATION AGAINST CODEC ATTACKS
Detection compares the selected score with a threshold τ, and it can augment hard matches with capped soft evidence from an auxiliary cluster affinity matrix RMARC: MULTI-BIT WATERMARKING FOR AUTOREGRESSIVE AUDIO GENERATION AGAINST CODEC ATTACKS
Decoding involves enumerating all 16 values at every symbol position to compute the hard symbol score Lhardq,a(s) = X t∈Oa jt=q log ηa1[ect = kt(s, ret)] + 1 − ηa KMARC: MULTI-BIT WATERMARKING FOR AUTOREGRESSIVE AUDIO GENERATION AGAINST CODEC ATTACKS
The final payload is initialized by a bit-wise majority voting of the n selected candidates, followed by a local refinement guided by pooled symbol scores and the pub-keyMARC: MULTI-BIT WATERMARKING FOR AUTOREGRESSIVE AUDIO GENERATION AGAINST CODEC ATTACKS
Robustness
MARC demonstrates robustness to watermark overwriting, showing that MARC preserves detection and recovery of the original payload when an additional watermark is embedded into the released audioMARC: MULTI-BIT WATERMARKING FOR AUTOREGRESSIVE AUDIO GENERATION AGAINST CODEC ATTACKS
The method maintains bit acc between 94.58% and 99.25% across all codec configurations evaluated over the three datasets, demonstrating the preservation of watermark evidence alongside accurate multi-bit recoveryMARC: MULTI-BIT WATERMARKING FOR AUTOREGRESSIVE AUDIO GENERATION AGAINST CODEC ATTACKS
Ablation Study
The full codec-union construction C0 achieves the highest bit acc of 98.17% and the lowest FAD of 0.2622 among the evaluated variants on clean audioMARC: MULTI-BIT WATERMARKING FOR AUTOREGRESSIVE AUDIO GENERATION AGAINST CODEC ATTACKS
The full decoder D0 improves both bit acc and TPR@FPR=1% over D1 and D2 under every evaluated codec attackMARC: MULTI-BIT WATERMARKING FOR AUTOREGRESSIVE AUDIO GENERATION AGAINST CODEC ATTACKS
The full configuration achieves a mean post-codec bit acc of 97.77%, compared with 97.24% without codec union, 94.71% with a single lap, and 96.
Improvements for AI systems
-
Codec-Aware Token Clustering for Robust Watermarking MARC integrates
intrinsic token representations with confusion patterns obtained through retokenization and multiple codecs to form a codec-aware token-cluster space,
allowing the system to model bothintrinsic token relationships and codec induced substitutions
simultaneously, which reduces the limitations of relying on a single source of information. -
Multi-Bit Payload Embedding The system enables
payload-driven cluster scheduling to embed a multi-bit watermark,
moving beyond existing methods that are limited tozero bit detection,
thereby allowing for the embedding and recovery of additional information into each generated audio sample. -
Robust Detection via Alignment Search MARC uses
alignment search + lap scoring
during extraction, where it evaluateswaveform variants, observation windows, and offsets of the watermark-step indices,
ensuring that even when token identities differ due to codec processing, the system can find evidence based onobserved clusters and their associated seeds.
-
Key-Conditioned Payload Recovery The system employs a verification protocol using a
public check value (pub-key) and a secret auxiliary suffix (pri-key),
where the pri-key is used forcandidate assessment
before the pub-key guidespost-voting refinement,
supporting recovery even when an attacker introduces an overwriting attack. -
Adaptation to Diverse Audio Attacks The framework demonstrates robustness against a wide array of attacks, including
neural audio codec attacks, signal processing, and desynchronization,
maintaining high bit extraction accuracy under diverse codec transformations and preserving the original payload during overwriting attacks.
Abstract
Generated audio is now used in a range of applications, creating a need to verify its origin after distribution and signal processing. This task is particularly challenging for autoregressive audio generation because codec processing can alter the token sequence recovered from the waveform. Such changes reduce the reliability of watermark detection and payload decoding. Existing methods construct token groups using either intrinsic token representations or substitution patterns caused by transformations. As a result, intrinsic token relationships and codec induced substitutions are modeled separately. In addition, most methods support only zero bit detection. They can determine whether a watermark is present but cannot distinguish individual generated outputs. We propose MARC, a multi-bit generative watermarking method for autoregressive audio generation. MARC integrates intrinsic token representations with confusion patterns obtained through retokenization and multiple codecs, forming a codec-aware token-cluster space. Within this space, payload-driven cluster scheduling is used to embed a multi-bit watermark, while detection and payload decoding are performed on retokenized observations. Experiments on speech, dialogue, and music generation show that MARC achieves an average of 97.3% bit extraction accuracy on unmodified watermarked audio and the watermark can still be extracted under diverse codec attacks. MARC also demonstrates robustness to overwriting attacks.
Sources
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs