Could LLM Watermark Detection be Public?

summary

Video file (mp4)

The gist

The gist The split-key construction and a novel, calibrated tampering test show that public detection carries a real but bounded liability, enabling transparency while keeping tampering with the

In short

The research investigates if public LLM watermarking detection is feasible, finding that it carries a real but bounded liability. A novel split-key construction allows for transparency while making tampering detectable. This method limits the threat of informed attacks, suggesting that publishing detectors can be beneficial for accountability.

Key concepts

Split-Key Construction
This defense uses two independent keys to route tokens to either a public or private channel with set probabilities. The detector scores these channels separately and then merges them into a fused score. This prevents the full watermark signal from being exposed privately, making it harder for attackers to tamper with the private signal.
Two-Stage Mechanism
The detection process involves two sequential tests. First, it checks if the fused score matches a threshold to determine if a watermark exists. Second, it runs a tampering test on that verdict. This second stage specifically looks for imbalances created by attackers trying to remove or forge the watermark.
Informed Attack
This attack scenario occurs when an attacker interacts with the detector to guide their edits at a lower distortion level. The split-key method is designed to counter this by routing tokens probabilistically, making it difficult for the attacker to precisely steer edits without being detected by the calibrated tampering test.
Public vs. Private Signal
The paper proposes releasing one of two watermark keys publicly while keeping the other private. This allows users to check for watermarks using a public detector, while the provider retains control over the private key, balancing transparency with security.

Terminology used across episodes

This episode discusses

The paper

Could LLM Watermark Detection be Public? · Read on arXiv

Georgios Milis, Tom Sander, Tomáš Souček, Heng Huang, Pierre Fernandez

FAIR · Meta Superintelligence Labs · University of Maryland

Watermarking large language models is popular for tracing chatbot and agentic outputs, yet detectors remain unreleased since exposing them could let attackers do targeted edits with the detector's feedback. However, watermarks are already vulnerable to uninformed tampering attacks. We thus first quantify whether a public detector would be an additional liability in a deployment setting at varying levels of access, from token-level scores to a binary verdict. Second, we introduce a split-key public-private watermarking method that exposes one key through a public detector while keeping the other for full verification and forensics. An informed attacker can only move the public signal, creating an imbalance between public and private scores. We introduce a statistical test for this imbalance, and combine it with the full key verdict in a two-stage mechanism. Third, we evaluate the split-key method on a wide range of removal and forgery attacks, comparing the uninformed to detector-informed settings. Public detection improves removal only at small edit budgets, since plain rephrasing already strips the watermark at a lower quality cost, but it does enable forgery, which the private pipeline can identify. Overall, releasing half of the watermark enables transparency and interoperability, and tampering with the released half stays detectable. This bounds the provider's liability and questions the need to keep detectors fully private.

Transcript

Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.

Nadia: Today's paper: "Could LLM Watermark Detection be Public?".

Elias: The gist The split-key construction and a novel, calibrated tampering test show that public detection carries a real but bounded liability,

Nadia: First, who's behind it and why it matters.

Paper summary: Nadia: So we're looking at this paper called "Could LLM Watermark Detection be Public?" It tackles the idea of whether making AI watermarking detectors public actually increases risk for the models and users.

Elias: Exactly. The main thrust here is that even if you make a detector public, it creates a real but bounded liability because of how an attacker can tamper with the released half of the key system while still having some recourse on the private side.

Priya: From my angle, I'm interested in what this means for actual privacy and measurement. The paper claims that this split-key construction exposes one key publicly while keeping another private for verification and forensics, which is a big deal because it lets you test if tampering is happening.

Nadia: Right. It’s not just about the existence of a detector, it's about the specific way they split those keys using Spub and Spriv to create two different p-values for testing null hypotheses.

Elias: That split benefits the platform because the private signal stays hidden, making it much harder for an attacker to tamper with that part, and that gap between public and private scores is itself informative about targeted tampering.

Priya: So what does this actually show us about how easy it is to strip a watermark or forge one when you have access to this detector?

Nadia: The paper focuses on two high-level attacks: rephrasing the whole text using global LLM paraphrasing, and local word edits like those you see with BERT attacks. They quantify the vulnerability this access adds by extending previous attacks into the informed setting in white-box and black-box tiers.

Elias: And then they introduce this two-stage mechanism that combines a watermark test on the fused score with a subsequent tampering test to check for removal or forgery across different nulls, like testing if the two channels are balanced.

Priya: What’s the actual finding here? Does this split-key method stop every kind of attack, or does it just make them harder to execute?

Nadia: The key finding is that this split limits the threat caused by informed attacks. Their strongest informed forgery succeeds on almost all carrier texts, but the majority of those attempts get caught. They also found that informed removal only helps at small edit budgets.

Elias: So the two-stage test they developed checks for different nulls: first if there is no watermark, and second if premoval or forgery has occurred by checking if the two channels are balanced against that null assumption of channel symmetry.

Priya: That sounds like a solid way to measure the actual data, rather than just relying on one simple pass/fail test. How does this apply to real-world deployment?

Nadia: The paper sets up their experiments using Qwen2 point 5-7B as the text generator, watermarked with TextSeal using two keys where alpha is set at zero point five and they use symmetric three gram contexts in each channel for the test.

Elias: They also extended this dual-key routing concept to other watermarks like Maryland and SynthID-Text, keeping the same setup of routing each position to a public key with probability of zero point five and the private key otherwise.

Priya: It’s interesting that they set the detection threshold for both the watermark test and the tampering test at an FPR of ten to the negative three, which is a pretty strict standard for what counts as a true false positive in this context.

Nadia: They also explicitly state their limitations: they mention that removal either rewrites text globally or edits it locally and often degrades quality, while forgery can piggyback on already-watermarked text or reverse-engineer the watermark from large corpora of watermarked outputs.

Elias: So what this paper ultimately suggests is that public detection carries a real but bounded liability because the split construction leaves removal close to what no detector allows, and keeps the forgery mostly detectable.

Priya: What does that mean for someone just listening to the show? It suggests that transparency can happen without completely opening up the system to any kind of exploitation.

Nadia: This work establishes that public detection carries a real but bounded liability, and it enables future research on watermark interoperability and transparency. The split construction offers a practical compromise between transparency and security because it keeps removal close to what the no-detector setting already allows, and keeps the forgery mostly detectable >

Conclusion: Nadia: So, basically, this paper is asking if making AI watermark detectors public actually creates more risk for the models and users involved.

Elias: It’s focused on a specific construction called a split-key system that lets you keep some keys private while releasing others publicly to test detection.

Priya: The core idea here is that this split allows the provider to test if someone is trying to tamper with their watermark signal without exposing the secret part of the verification process.

Nadia: Exactly, and they show how this setup creates a real but bounded liability for publishing a detector because an attacker can still do things on the private side.

Elias: They used two separate keys, Spub and Spriv, and they score them separately before fusing them to get one final verdict for the platform.

Priya: What’s interesting is that this gap between the public score and the private score itself becomes a piece of information you can use to spot if someone is trying to manipulate things.

Nadia: So, even though detection is public, it’s not a total open door for attackers; it just creates a measurable risk they have to manage.

Elias: The two-stage mechanism they built tests the watermark on that fused score and then immediately checks if any removal or forgery happened using another test based on channel symmetry.

Priya: The results show that this split construction limits how much harm an informed attacker can do, meaning they can't completely strip the watermark easily.

Nadia: They also found that forgeries are mostly detectable with this method, which is a good sign for accountability when you publish detection tools.

Elias: This work suggests that a dual-key routing approach is a practical way to balance transparency with the security needed to keep watermarking effective.

Priya: It opens up questions about how we can build systems that allow public oversight without giving away all the secrets needed for tampering.

More episodes

← Home