Could LLM Watermark Detection be Public?
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.
Nadia: Today's paper: "Could LLM Watermark Detection be Public?".
Elias: The gist The split-key construction and a novel, calibrated tampering test show that public detection carries a real but bounded liability,
Nadia: First, who's behind it and why it matters.
Paper summary: Nadia: So we're looking at this paper called "Could LLM Watermark Detection be Public?" It tackles the idea of whether making AI watermarking detectors public actually increases risk for the models and users.
Elias: Exactly. The main thrust here is that even if you make a detector public, it creates a real but bounded liability because of how an attacker can tamper with the released half of the key system while still having some recourse on the private side.
Priya: From my angle, I'm interested in what this means for actual privacy and measurement. The paper claims that this split-key construction exposes one key publicly while keeping another private for verification and forensics, which is a big deal because it lets you test if tampering is happening.
Nadia: Right. It’s not just about the existence of a detector, it's about the specific way they split those keys using Spub and Spriv to create two different p-values for testing null hypotheses.
Elias: That split benefits the platform because the private signal stays hidden, making it much harder for an attacker to tamper with that part, and that gap between public and private scores is itself informative about targeted tampering.
Priya: So what does this actually show us about how easy it is to strip a watermark or forge one when you have access to this detector?
Nadia: The paper focuses on two high-level attacks: rephrasing the whole text using global LLM paraphrasing, and local word edits like those you see with BERT attacks. They quantify the vulnerability this access adds by extending previous attacks into the informed setting in white-box and black-box tiers.
Elias: And then they introduce this two-stage mechanism that combines a watermark test on the fused score with a subsequent tampering test to check for removal or forgery across different nulls, like testing if the two channels are balanced.
Priya: What’s the actual finding here? Does this split-key method stop every kind of attack, or does it just make them harder to execute?
Nadia: The key finding is that this split limits the threat caused by informed attacks. Their strongest informed forgery succeeds on almost all carrier texts, but the majority of those attempts get caught. They also found that informed removal only helps at small edit budgets.
Elias: So the two-stage test they developed checks for different nulls: first if there is no watermark, and second if premoval or forgery has occurred by checking if the two channels are balanced against that null assumption of channel symmetry.
Priya: That sounds like a solid way to measure the actual data, rather than just relying on one simple pass/fail test. How does this apply to real-world deployment?
Nadia: The paper sets up their experiments using Qwen2 point 5-7B as the text generator, watermarked with TextSeal using two keys where alpha is set at zero point five and they use symmetric three gram contexts in each channel for the test.
Elias: They also extended this dual-key routing concept to other watermarks like Maryland and SynthID-Text, keeping the same setup of routing each position to a public key with probability of zero point five and the private key otherwise.
Priya: It’s interesting that they set the detection threshold for both the watermark test and the tampering test at an FPR of ten to the negative three, which is a pretty strict standard for what counts as a true false positive in this context.
Nadia: They also explicitly state their limitations: they mention that removal either rewrites text globally or edits it locally and often degrades quality, while forgery can piggyback on already-watermarked text or reverse-engineer the watermark from large corpora of watermarked outputs.
Elias: So what this paper ultimately suggests is that public detection carries a real but bounded liability because the split construction leaves removal close to what no detector allows, and keeps the forgery mostly detectable.
Priya: What does that mean for someone just listening to the show? It suggests that transparency can happen without completely opening up the system to any kind of exploitation.
Nadia: This work establishes that public detection carries a real but bounded liability, and it enables future research on watermark interoperability and transparency. The split construction offers a practical compromise between transparency and security because it keeps removal close to what the no-detector setting already allows, and keeps the forgery mostly detectable >
Conclusion: Nadia: So, basically, this paper is asking if making AI watermark detectors public actually creates more risk for the models and users involved.
Elias: It’s focused on a specific construction called a split-key system that lets you keep some keys private while releasing others publicly to test detection.
Priya: The core idea here is that this split allows the provider to test if someone is trying to tamper with their watermark signal without exposing the secret part of the verification process.
Nadia: Exactly, and they show how this setup creates a real but bounded liability for publishing a detector because an attacker can still do things on the private side.
Elias: They used two separate keys, Spub and Spriv, and they score them separately before fusing them to get one final verdict for the platform.
Priya: What’s interesting is that this gap between the public score and the private score itself becomes a piece of information you can use to spot if someone is trying to manipulate things.
Nadia: So, even though detection is public, it’s not a total open door for attackers; it just creates a measurable risk they have to manage.
Elias: The two-stage mechanism they built tests the watermark on that fused score and then immediately checks if any removal or forgery happened using another test based on channel symmetry.
Priya: The results show that this split construction limits how much harm an informed attacker can do, meaning they can't completely strip the watermark easily.
Nadia: They also found that forgeries are mostly detectable with this method, which is a good sign for accountability when you publish detection tools.
Elias: This work suggests that a dual-key routing approach is a practical way to balance transparency with the security needed to keep watermarking effective.
Priya: It opens up questions about how we can build systems that allow public oversight without giving away all the secrets needed for tampering.
Georgios Milis, Tom Sander, Tomáš Souček, Heng Huang, Pierre Fernandez
FAIR · Meta Superintelligence Labs · University of Maryland
cs.CR, cs.LG
Submitted: 2026-10-08
Updated: 2026-10-08
Code: https://github.com/facebookresearch/textseal
License: http://creativecommons.org/licenses/by/4.0/
The gist: The gist The split-key construction and a novel, calibrated tampering test show that public detection carries a real but bounded liability, enabling transparency while keeping tampering with the
Key concepts
- Split-Key Construction
- This defense uses two independent keys to route tokens to either a public or private channel with set probabilities. The detector scores these channels separately and then merges them into a fused score. This prevents the full watermark signal from being exposed privately, making it harder for attackers to tamper with the private signal.
- Two-Stage Mechanism
- The detection process involves two sequential tests. First, it checks if the fused score matches a threshold to determine if a watermark exists. Second, it runs a tampering test on that verdict. This second stage specifically looks for imbalances created by attackers trying to remove or forge the watermark.
- Informed Attack
- This attack scenario occurs when an attacker interacts with the detector to guide their edits at a lower distortion level. The split-key method is designed to counter this by routing tokens probabilistically, making it difficult for the attacker to precisely steer edits without being detected by the calibrated tampering test.
- Public vs. Private Signal
- The paper proposes releasing one of two watermark keys publicly while keeping the other private. This allows users to check for watermarks using a public detector, while the provider retains control over the private key, balancing transparency with security.
Terminology
Summary
The gist The split-key construction and a novel, calibrated tampering test show that public detection carries a real but bounded liability, enabling transparency while keeping tampering with the released half detectable
Background and Threat Model
LLM watermarking alters the tokens’ sampling in a key-dependent way to leave a statistical trace only the key holder can test for The attacker’s goal is either to remove the watermark so that AI-generated text evades detection, or to forge it, i.e. embed it into text the platform’s model did not generate The threat model considers four scenarios of increasing detector transparency: uninformed no-box, black-box, gray-box, and white-box White-box is the most permissive setting so it upper-bounds the liability of publishing a detector
Attacks on a Public Detector
Tampering with LLM watermarks can be done by two high-level operations: Rephrasing rewrites the whole text through global LLM paraphrasing or back-translation, whereas word edits apply local lexical substitutions, typically via BERT-based attacks The paper quantifies the vulnerability this access adds by extending previous attacks to the informed setting in white-box and black-box tiers Forgery can be done by adding watermarking signal to falsely blame a provider, for instance by inserting a name into a defamatory sentence or fabricating compliance evidence
Split-Key as a Defense Against Informed Attacks
The split-key construction involves scoring against two independent keys, Spub and Spriv, each tested against its own null to give two p-values ppub and ppriv The public detector returns ppub to any user, while the platform keeps the detector based on the fused score, which tests αSpub + (1 − α)Spriv against the full null for its forensic verdict pfus This split benefits the provider on two counts: first, the private signal is never exposed and is thus much harder to tamper with second, the resulting gap between the public and private scores is itself informative, which we exploit in the next paragraph to detect targeted tampering
Two-Stage Mechanism
The two-stage mechanism consists of the watermark test on the fused score followed by the tampering test, similar to an image watermarking construction of Evennou and Kijak The first stage compares pfus to the threshold set by the target False Positive Rate (FPR), which decides the binary watermark label yfus The second stage re-evaluates that watermark verdict with the tampering test This tests different nulls: pfus that the text carries no watermark, and premoval or pforgery that its two channels are balanced
Conclusion
The work establishes that public detection carries a real but bounded liability, and enable future research on watermark interoperability and transparency A split construction can be used as a practical compromise between transparency and security, since it leaves removal close to what the no-detector setting already allows, and keeps the forgery it enables mostly detectable Watermark detection could therefore be public: even the strongest attacks that a transparent detector enables leave the provider with a liability that is real but bounded by the key split >
How it works
The paper introduces an informed attack where an attacker queries the detector to steer edits at lower distortion The split-key method routes each token to one of two keys with probabilities α and 1 − α, and scores are merged into a fused score S = αS1 + (1 − α)S2 A novel, calibrated tampering test identifies the public-private imbalance an informed attacker creates
Key Findings
The split limits the threat caused by informed attacks, as our strongest informed forgery succeeds on almost all carrier texts yet the majority is caught, while informed removal helps only at small edit budgets The two-stage test tests different nulls: pfus that the text carries no watermark, and premoval or pforgery that its two channels are balanced The tampering test is calibrated on its null assumption of channel symmetry
Experimental Setup
Experiments use Qwen2.5-7B as the text generator, watermarked with TextSeal using 2 keys, α=0.5 and symmetric h=3-gram contexts in each channel The detection threshold for both the watermark test and the tampering test is set at FPR of 10−3 Attack parameters include rephrasing via autoregressive regeneration and in-place word edits steered by the detector
Extension to Other Schemes
Dual-key routing extends to other watermarks like Maryland and SynthID-Text under the same construction of routing each position to the public key with probability α=0.5, and to the private key otherwise Both schemes use the same context of h=3 tokens per channel
Ethics Statement
The work studies whether LLM watermark detection can be public, and we introduce a way to publish detection results: the provider releases one of two watermark keys through a public detector and keeps the other private We therefore expect the overall impact of this work to be positive >
Reproducibility Statement
Experiments build mainly on https://github.com/facebookresearch/textseal, which has a permissive license (Apache-2.0) and all models and datasets are publicly available We will release the experiment code with a permissive license for the camera-ready version >
References
Scott Aaronson. Watermarking of large language models. Talk, Simons Institute for the Theory of Computing, 2022 Anthropic. How Claude marks AI-generated content, 2026 Toluwani Aremu et al. Mitigating watermark forgery in generative models via randomized key selection, arXiv preprint arXiv:2507.07871, 2025 Shane Arora et al. Calmqa: Exploring culturally specific long-form question answering across 23 languages, In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 11772–11817, 2025 Shariq Bashir. Asymmetric watermarking for large language models with public and private verification, IEEE Access, 2026 California State Legislature. SB-942 California AI Transparency Act: Cal. Bus. & Prof. Code § 22757 – disclosure and AI detection tools, Enacted Sept. 19, 2024, effective Jan. 1, 2026, 2024 Ruibo Chen et al. De-mark: Watermark removal in large language models, In International Conference on Machine Learning, pages 9316–9333. PMLR, 2025 Yixin Cheng et al. Revealing weaknesses in text watermarking through self-information rewrite attacks, In International Conference on Machine Learning, pages 9982–10009. PMLR, 2025 Miranda Christ and Sam Gunn. Pseudorandom error-correcting codes, In Annual International Cryptology Conference, pages 325–347. Springer, 2024 John Kirchenbauer et al. A watermark for large language models, In International conference on machine learning, pages 17061–17084.
Improvements for AI systems
-
Bold key split-key watermarking construction: Alice should
split the watermark signal across two keys, releasing a detector that uses only the public half kpub while keeping kpriv private,
enabling transparency without exposing the full secret to an attacker. -
Two-stage mechanism for detection: Implement a
two-stage mechanism, consisting of the watermark test on the fused score, followed by the tampering test,
which allows Alice to determine if a text iswatermarked or not
before performing a more rigorous check for removal or forgery. -
Calibrated tampering test: Introduce a
novel, calibrated tampering test that identifies the public-private imbalance an informed attacker creates,
which tests the channel discrepancy by calculating the score difference T, and uses this to determine if a text isremoved or forged.
-
Adaptive attack resilience: Design removal attacks to use an
overshoot factor m
when stopping at a verdict flip, allowing attackers to exploit the system's stopping rule while still being caught by the two-stage test for strong edits. -
Context-aware routing: Use
unequal h-gram contexts
where thewide public context decays faster, so the gap grows with editing,
ensuring that an attacker cannot maintain a significant score imbalance when tampering occurs.
Abstract
Watermarking large language models is popular for tracing chatbot and agentic outputs, yet detectors remain unreleased since exposing them could let attackers do targeted edits with the detector's feedback. However, watermarks are already vulnerable to uninformed tampering attacks. We thus first quantify whether a public detector would be an additional liability in a deployment setting at varying levels of access, from token-level scores to a binary verdict. Second, we introduce a split-key public-private watermarking method that exposes one key through a public detector while keeping the other for full verification and forensics. An informed attacker can only move the public signal, creating an imbalance between public and private scores. We introduce a statistical test for this imbalance, and combine it with the full key verdict in a two-stage mechanism. Third, we evaluate the split-key method on a wide range of removal and forgery attacks, comparing the uninformed to detector-informed settings. Public detection improves removal only at small edit budgets, since plain rephrasing already strips the watermark at a lower quality cost, but it does enable forgery, which the private pipeline can identify. Overall, releasing half of the watermark enables transparency and interoperability, and tampering with the released half stays detectable. This bounds the provider's liability and questions the need to keep detectors fully private.
Sources
- Mitigating Watermark Forgery in Generative Models via Randomized Key Selection
- Gemma 3 Technical Report
- RLSpoofer: A Lightweight Evaluator for LLM Watermark Spoofing Resilience
- RoBERTa: A Robustly Optimized BERT Pretraining Approach
- Qwen2.5 Technical Report
- TextSeal: A Localized LLM Watermark for Provenance & Distillation Protection
- Permute-and-Flip: An optimally stable and watermarkable decoder for LLMs
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs