Watermark Forensics for Generative Models: An Information-Theoretic Perspective
Xiaoyu Li, Zheng Gao, Xiaoyan Feng, Jiaojiao Jiang, Yulei Sui, Jiankun Hu
cs.CR, cs.IT, cs.LG
Submitted: 2026-07-14
Comments: The abstract has been shortened to comply with arXiv's length limit
License: http://creativecommons.org/licenses/by/4.0/
The gist: A watermark in a generative model's output is usually asked only whether a text is machine-made.
Terminology
Abstract
A watermark in a generative model's output is usually asked only whether a text is machine-made. The same mark can do more: attribute it to the user who produced it, extract a hidden payload, or localize the part that survives editing. These form a forensic ladder, and we ask what each rung costs in the sample length n. One object organizes the answers. Let S be the secret the mark carries (a user's identity or payload), and let the information profile nu(t)=I(S;X t X<t) record how much the t-th token reveals about S given the earlier ones. Its total mass pays for attribution and extraction; how that mass is spread pays for localization; and detection alone is paid for not by information but by presence, the distance from the marked to the unmarked distribution. The literature's two quality models, a mark subtle on every token and one that stamps a few tokens loudly, are two incomparable ways of capping this profile. Our main theorem settles the ladder's entropy column. For statistically distortion-free schemes, attributing a text to one of N users costs (N/h) tokens over every stationary-ergodic source of entropy rate h, sharp to a (1+o(1)) factor: to our knowledge the first tight entropy-rate law for multi-user attribution (via exact alignment). The natural collision-counting analysis overcharges without bound; only a decoder thresholding each candidate by its own realized surprisal attains the rate while almost never implicating an innocent user. A matching converse makes the law two-sided, and extraction of an-bit payload costs (/h). Two gaps are real, not modeling artifacts: a (N) -token window in which a text is provably machine-made yet unattributable, and a footprint-resolution uncertainty principle. Experiments on GPT-2, Pythia-410M, and Qwen2.5 recover the predicted constants.
Sources
- LLM Watermarking Using Mixtures and Statistical-to-Computational Gaps
- PASA: A Principled Embedding-Space Watermarking Approach for LLM-Generated Text under Semantic-Invariant Attacks
- SEAL: Semantic Aware Image Watermarking
- SHIFT: Stochastic Hidden-Trajectory Deflection for Removing Diffusion-based Watermark
- Fast segmentation of watermarked texts from large language models through an epidemic change-point framework
- Towards Better Statistical Understanding of Watermarking LLMs
- Sample Complexities of Estimating Gumbel--Max Watermark Proportions with and without Reduction to Pivotal Statistics
- Rotation-Invariant Spherical Watermarking via Third-Order SO(3) Representation Coupling
- Pseudorandom Error-Correcting Codes
- The Stable Signature: Rooting Watermarks in Latent Diffusion Models
- Multi-use LLM Watermarking and the False Detection Problem
- TRACE: A Two-Channel Robust Attribution Watermark via Complementary Embeddings for LLM-Agent Trajectories
- SLICE: Semantic Latent Injection via Compartmentalized Embedding for Image Watermarking
- ArcMark: Distortion-Free Multi-Byte LLM Watermark via Optimal Transport
- A Unified Framework for LLM Watermarks
- Covert Multi-bit LLM Watermarking: An Information Theory and Coding Approach
- Distributional Information Embedding: A Framework for Multi-bit Watermarking
- Fundamental Trade-Offs in Multi-Bit Watermarking of Stochastic Processes
- Optimal Watermark Generation under Type I and Type II Errors
- Towards Optimal Statistical Watermarking
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs