Watermarking Diffusion Language Models
cs.LG, cs.AI, cs.CR
Submitted: 2025-09-29
Updated: 2026-09-17
Code: https://github.com/open-compass/opencompass
License: http://creativecommons.org/licenses/by/4.0/
The gist: We introduce the first watermark tailored for diffusion language models (DLMs), an emergent LLM paradigm able to generate tokens in arbitrary order, in contrast to standard autoregressive language
Terminology
Abstract
We introduce the first watermark tailored for diffusion language models (DLMs), an emergent LLM paradigm able to generate tokens in arbitrary order, in contrast to standard autoregressive language models (ARLMs) which generate tokens sequentially. While there has been much work in ARLM watermarking, a key challenge when attempting to apply these schemes directly to the DLM setting is that they rely on previously generated tokens, which are not always available with DLM generation. In this work we address this challenge by: (i) applying the watermark in expectation over the context even when some context tokens are yet to be determined, and (ii) promoting tokens which increase the watermark strength when used as context for other tokens. This is accomplished while keeping the watermark detector unchanged. Our experimental evaluation demonstrates that the DLM watermark leads to a >99% true positive rate with minimal quality impact and achieves similar robustness to existing ARLM watermarks, enabling for the first time reliable DLM watermarking.
Sources
- Block Diffusion: Interpolating Between Autoregressive and Diffusion Language Models
- Program Synthesis with Large Language Models
- The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Training Verifiers to Solve Math Word Problems
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
- AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback
- Towards Watermarking of Open-Source LLMs
- Better & Faster Large Language Models via Multi-token Prediction
- The Llama 3 Herd of Models
- Train for the Worst, Plan for the Best: Understanding Token Ordering in Masked Diffusions
- Robust Distortion-free Watermarks for Language Models
- Mercury: Ultra-Fast Language Models Based on Diffusion
- Jailbreaking Prompt Attack: A Controllable Adversarial Attack against Diffusion Models
- Large Language Diffusion Models
- GPT-4 Technical Report
- Your Absorbing Discrete Diffusion Secretly Models the Conditional Distributions of Clean Data
- MarkLLM: An Open-Source Toolkit for LLM Watermarking
- Denoising Diffusion Implicit Models
- Tree-Ring Watermarks: Fingerprints for Diffusion Images that are Invisible and Robust
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks